fleet01 as-built: the full second-host setup, and the gaps a fleet manager would have to own #156

Open
opened 2026-08-23 14:15:46 +02:00 by ltms · 0 comments
Owner

Reference ticket. fleet01 is our second fleet host and the first one that is not the operator's Mac. Everything below was measured on the host on 2026-08-23, not recalled. It is written down because we are starting to talk about a fleet manager, and the useful input for that talk is not the design we want — it is the set of things a human had to do by hand to make a second host work, and the things that still have no owner.

Where I did not check something myself, I say so.


1. The host

name / OS fleet01, Ubuntu 24.04.1 LTS, kernel 6.8.0-45
size 8 CPU, 11 GiB RAM
network 10.10.20.13/24 on ens18; 172.17.0.1 on docker0
user ltms (uid 1000), in sudo and docker
login shell /usr/bin/zsh
Java / Maven OpenJDK 25.0.3, Maven 3.8.7
herdr 0.8.0, protocol 19

The network is one-way. The Mac reaches fleet01. fleet01 cannot reach the Mac. This one fact decided more of the design than anything else, and any fleet manager has to assume it is normal rather than a mistake.

There is no GUI and no display. Everything happens over ssh.

2. Starting herdr headless — the part that was hardest

A plain herdr needs an interactive terminal to stay alive, so it dies when the ssh session ends. The working start is fleetd-run/start-herdr.sh. It does three things that are not obvious:

  1. It runs herdr as a named session (herdr --session fleet01), not the default one.
  2. It wraps it in setsid script -qfec … /dev/null, so herdr gets a real pty and survives the ssh logout.
  3. Inside that pty it runs stty rows 50 cols 200 before exec'ing herdr.

Step 3 is the one that cost real time. herdr sizes each new pane from the attached client's view. Under a pty with no window size the client reports 0x0. Every pane.split and workspace.create then fails with ghostty error -2, because libghostty refuses a 0x0 surface. The daemon reports this three steps later as a plain spawn failure, so the message you see never mentions the size.

Two herdr sessions now run on this host at the same time: the default one (socket ~/.config/herdr/herdr.sock) and fleet01 (socket ~/.config/herdr/sessions/fleet01/herdr.sock). fleetd is pinned to the second one with herdrSocket:. A bare herdr … command on the CLI talks to the first one.

I got this wrong myself during this work. I ran herdr pane list, saw no lead, and told the operator that the fleet01 lead was missing. It was there the whole time, in the other session. I asked for a restart that was not needed. Any tool that shows fleet state must say which session it read, or it will produce confident wrong answers.

The default session still holds two panes sitting in ~/LTMS/kb, one running claude. Nothing owns or reaps them.

3. Running fleetd

Config file is bridged/fleetd.yaml. Note the name: the Mac uses bridged.yaml. Both hosts also carry a fleetd.example.yaml, which makes guessing the wrong name easy.

Restart is fleetd-run/restart.sh. It is small and deliberately does four things:

  • starts java from a login shell (setsid zsh -lc "exec java -jar target/bridged.jar fleetd.yaml")
  • waits for the old process to actually exit
  • polls /healthz
  • proves the pid changed

Nothing supervises either process. I checked for user-level systemd units, system-level units, and cron entries. There are none. The Mac has launchd plus scripts/bridged-launchd-wrapper.sh; deploy/bridged.service exists in the repo but is not installed here, and it has the login-shell secret defect noted in the wiki. So on fleet01 a reboot, an OOM kill, or a crash leaves the host with no fleet and nothing that notices.

Deployment is fully manual and per host: git pull, mvn clean package, restart.sh. There is no build artefact shared between hosts. Today the Mac and fleet01 are both on 3f4ac2b, but only because I did each one by hand.

4. Config: profiles and roles

Four profiles, all reaching models through llm.ltms.dev:

profile kind model weight maxLoad
gx opencode gx/deepseek-v4-flash 100 2
xf opencode opencode/x-preview-f-free 80 5
local claude-code deepseek-v4-flash via /anthropic 10 2
opus claude-code claude-opus-5, subscription: true 0 1

Roles: leaders = opus; developers = xf and gx; reviewers = gx. Placement is weighted. Lifecycle: idleTtlSeconds: 1800, contextCap: 8, drainTimeoutSeconds: 120.

The lead slot is pinned by tab label lead: opus, with cwd: /home/ltms/LTMS/kb. That matches the operator's rule that a lead is started in the project the fleet works on.

The reviewer separation is weak and the config says so. With only two usable backends, gx reviews diffs that gx may have written. That is a real limit of a small host, not an oversight.

Two config comments are now wrong, and both are the kind of drift a fleet manager would have to prevent:

  • The fleet.leaders block says the opus profile is "NOT defined above on purpose". It is defined above.
  • The same block says a lead cannot start until someone completes claude /login on this host. That login now exists — ~/.claude.json has an oauthAccount. Whether an opus lead can actually spawn is a separate question, because #140 says every subscription: true claude-code profile fails to spawn. I did not re-test #140 on this host.

5. Credentials — better here than on the Mac

fleet01 holds only four values, in ~/.fleet/secrets.sh (mode 600): AI_GATEWAY_TOKEN, WORKER_GITEA_TOKEN, GITEA_HOST, LAVINMQ_URI. The Mac's store holds dozens, including an AWS AdministratorAccess key.

The file is sourced from ~/.zprofile, which only a login shell reads. A herdr member pane on Linux is an interactive but non-login shell — measured as a bare /usr/bin/zsh in /proc/<pid>/cmdline. So members never inherit these values at all. The CB-633 scrub is a second line of defence here rather than the only one. This is the better arrangement and it is worth copying back to the Mac.

fleetd.yaml contains no literal secret. Every credential is referenced by env var name (tokenEnv, gitTokenEnv, and now broker.uriEnv).

But the git remotes leak the forge token

~/LTMS/kb/.git/config has the token written into the remote URL:

https://<WORKER_GITEA_TOKEN>@git.ltms.dev/akb/kb.git

Both worker worktrees under ~/LTMS/.fleet-worktrees/ inherit the same URL. The file is mode 664.

This defeats the credential scrub completely. The scrub removes environment variables; it cannot remove a token written into a file inside the very repo the member is told to work in. A member that never sees WORKER_GITEA_TOKEN in its environment can still read it with git remote -v.

I found this by accident, and in doing so I printed the token into an operator transcript. It needs rotating, and so does AI_GATEWAY_TOKEN, which I exposed separately in the same session by using ${VAR:-default} to test whether it was set — that expands to the value.

Both mistakes are the same shape and both are worth designing against: the credential leaves through a channel nobody classified as a credential channel. A fleet manager that reports host state will print git remotes, process lists and config files. Every one of those is a leak path.

Suggested immediate fix, separate from any manager work: use ssh remotes, or a credential helper, and never git clone with the token inline.

A false alarm that teaches operators to ignore warnings

Startup logs this on fleet01:

startup secret BRIDGED_WORKER_TOKEN: MISSING (profile 'xf' tokenEnv)

xf has no tokenEnv in the config. The name comes from FleetConfig.Profile's compact constructor, which defaults tokenEnv to "BRIDGED_WORKER_TOKEN" for every profile that leaves it unset — including opencode profiles that need no token at all. So the daemon warns, at every boot, about a variable that nobody configured and nothing needs.

The report itself is good and I want to keep it. The default is what makes it lie. A profile that needs no token should produce no line.

6. The shared broker

One LavinMQ 2.9.1 container runs on fleet01: fleet-lavinmq, image cloudamqp/lavinmq:2.9.1, volume fleet-lavinmq-data, --restart unless-stopped, listening on 0.0.0.0:5672 and 0.0.0.0:15672.

fleet vhost user
Mac /mac fleet-mac
fleet01 /fleet01 fleet-fleet01

guest is administrator on / only, and was removed from both fleet vhosts. Neither fleet user has a management tag, so both get 401 from the HTTP API. Inspect the broker with guest:guest from on fleet01.

The broker had to live here, not on the Mac. Two measured facts forced it: fleet01 cannot reach the Mac, and — until #152 landed today — an unreachable broker at boot stopped the daemon from starting at all. The second is fixed; the first is not.

Isolation is enforced, not just agreed. I ran the AMQP contract suite twice with the Mac's credential: 8/8 pass against /mac, and 8 errors out of 8 against /fleet01. The permission boundary is real.

Related, and now more likely to matter: #154 — the inbox pulls whole queues into an unbounded in-memory map, so the broker-side limits never apply.

7. Project state

~/LTMS/kb is the akb/kb checkout the lead works in. It sits on branch llm-backend-configurable, 3 commits behind origin/main, with a modified vendor/graphiti submodule. Two worker worktrees are left over from earlier runs, both dirty, both on branches that were already merged:

  • .fleet-worktrees/57b308-4 → worker/debrand-code-assumptions-b-57ac39-4
  • .fleet-worktrees/fec975-1 → worker/debrand-identity-587928-1

Nothing reaps these. lifecycle.idleTtlSeconds reaps members, not their worktrees.


8. What a fleet manager would have to own

This is the point of the ticket. Everything above is one host, set up by hand. Ordered by how much it hurt.

1. Nothing keeps anything alive. No supervisor for fleetd, none for herdr. A reboot ends the fleet silently. The Mac solved this with launchd; the Linux answer exists in the repo as deploy/bridged.service but is not installed and carries a known secret defect. This is the single biggest gap.

2. A dead lead is never restored, and nothing says so. LeadLauncher.ensureLeads() runs once, inline in Fleetd.main(). There is no supervision loop. If the lead pane dies an hour later, the slot stays empty until someone restarts the daemon, and no log line reports it. Related: #150 — leads always report ready: false, so the status endpoint cannot be used to detect this either.

3. Deployment is manual, per host, and drifts. git pull + mvn package + restart.sh, typed by a human on each machine. There is no shared artefact and no record of which host runs which commit. Today both hosts match only because I checked.

4. Config drift has no guard. Different file names per host (bridged.yaml vs fleetd.yaml), different profile sets, and comments that are already wrong. Config files are gitignored, so a change to what a key means can ship green and break a live host.

5. There is no cross-host anything. Two fleets now share a broker, and that is all they share. fleet_list shows only local leads. A Mac lead cannot see, message, or coordinate with a fleet01 lead. Logs are local files on each host. If a fleet manager is meant to answer "what is my whole fleet doing", none of that exists yet.

6. Diagnosis reads the wrong source by default. Two herdr sessions on one host, and the plain CLI talks to the wrong one. Any status a manager reports must name the socket it read.

7. Credentials leak through channels nobody classified. Git remote URLs today. Process lists and config dumps tomorrow. A manager that displays host state is a new leak surface by construction, and it should be designed with redaction as a rule rather than as care.

8. Nothing cleans up. Stale worktrees, stale panes in the other herdr session, 173 orphan queues left on the old Mac-local broker. Each is small. Together they are the reason a host slowly stops being understandable.

What already works and should not be redesigned

  • The login-shell rule for starting the daemon. It is the difference between working credentials and a silent failure that appears hours later.
  • Keeping secrets in ~/.zprofile so non-login member panes never inherit them. Defence before the scrub, not instead of it.
  • Naming env vars in config instead of values. With broker.uriEnv there is now no literal secret in either host's config.
  • The --check shape of scripts/redeploy-bridged.sh: read-only, reports whether each named variable resolves, never prints a value. That is the right model for anything a manager runs against a host.

Not verified

  • I did not re-test #140 (subscription profiles failing to spawn) on this host, so I cannot say whether an opus lead can start now that the login exists.
  • I did not launch a member and watch a queue appear on /fleet01; the broker connection is proven, an end-to-end delegation on this host is not.
  • I did not measure whether the two leftover worktrees still hold work worth keeping.
Reference ticket. `fleet01` is our second fleet host and the first one that is not the operator's Mac. Everything below was **measured on the host on 2026-08-23**, not recalled. It is written down because we are starting to talk about a *fleet manager*, and the useful input for that talk is not the design we want — it is the set of things a human had to do by hand to make a second host work, and the things that still have no owner. Where I did not check something myself, I say so. --- # 1. The host | | | |---|---| | name / OS | `fleet01`, Ubuntu 24.04.1 LTS, kernel 6.8.0-45 | | size | 8 CPU, 11 GiB RAM | | network | `10.10.20.13/24` on `ens18`; `172.17.0.1` on `docker0` | | user | `ltms` (uid 1000), in `sudo` and `docker` | | login shell | `/usr/bin/zsh` | | Java / Maven | OpenJDK 25.0.3, Maven 3.8.7 | | herdr | 0.8.0, protocol 19 | **The network is one-way.** The Mac reaches `fleet01`. `fleet01` cannot reach the Mac. This one fact decided more of the design than anything else, and any fleet manager has to assume it is normal rather than a mistake. There is no GUI and no display. Everything happens over ssh. # 2. Starting herdr headless — the part that was hardest A plain `herdr` needs an interactive terminal to stay alive, so it dies when the ssh session ends. The working start is `fleetd-run/start-herdr.sh`. It does three things that are not obvious: 1. It runs herdr as a **named session** (`herdr --session fleet01`), not the default one. 2. It wraps it in `setsid script -qfec … /dev/null`, so herdr gets a real pty and survives the ssh logout. 3. Inside that pty it runs `stty rows 50 cols 200` **before** exec'ing herdr. Step 3 is the one that cost real time. herdr sizes each new pane from the attached client's view. Under a pty with no window size the client reports 0x0. Every `pane.split` and `workspace.create` then fails with `ghostty error -2`, because libghostty refuses a 0x0 surface. The daemon reports this three steps later as a plain spawn failure, so the message you see never mentions the size. **Two herdr sessions now run on this host at the same time**: the default one (socket `~/.config/herdr/herdr.sock`) and `fleet01` (socket `~/.config/herdr/sessions/fleet01/herdr.sock`). `fleetd` is pinned to the second one with `herdrSocket:`. A bare `herdr …` command on the CLI talks to the **first** one. I got this wrong myself during this work. I ran `herdr pane list`, saw no lead, and told the operator that the fleet01 lead was missing. It was there the whole time, in the other session. I asked for a restart that was not needed. Any tool that shows fleet state must say **which session** it read, or it will produce confident wrong answers. The default session still holds two panes sitting in `~/LTMS/kb`, one running `claude`. Nothing owns or reaps them. # 3. Running fleetd Config file is `bridged/fleetd.yaml`. Note the name: the **Mac** uses `bridged.yaml`. Both hosts also carry a `fleetd.example.yaml`, which makes guessing the wrong name easy. Restart is `fleetd-run/restart.sh`. It is small and deliberately does four things: - starts java from a **login shell** (`setsid zsh -lc "exec java -jar target/bridged.jar fleetd.yaml"`) - waits for the old process to actually exit - polls `/healthz` - proves the pid changed **Nothing supervises either process.** I checked for user-level systemd units, system-level units, and cron entries. There are none. The Mac has launchd plus `scripts/bridged-launchd-wrapper.sh`; `deploy/bridged.service` exists in the repo but is not installed here, and it has the login-shell secret defect noted in the wiki. So on `fleet01` a reboot, an OOM kill, or a crash leaves the host with no fleet and nothing that notices. Deployment is fully manual and per host: `git pull`, `mvn clean package`, `restart.sh`. There is no build artefact shared between hosts. Today the Mac and `fleet01` are both on `3f4ac2b`, but only because I did each one by hand. # 4. Config: profiles and roles Four profiles, all reaching models through `llm.ltms.dev`: | profile | kind | model | weight | maxLoad | |---|---|---|---|---| | `gx` | opencode | `gx/deepseek-v4-flash` | 100 | 2 | | `xf` | opencode | `opencode/x-preview-f-free` | 80 | 5 | | `local` | claude-code | `deepseek-v4-flash` via `/anthropic` | 10 | 2 | | `opus` | claude-code | `claude-opus-5`, `subscription: true` | 0 | 1 | Roles: leaders = `opus`; developers = `xf` and `gx`; reviewers = `gx`. Placement is `weighted`. Lifecycle: `idleTtlSeconds: 1800`, `contextCap: 8`, `drainTimeoutSeconds: 120`. The lead slot is pinned by tab label `lead: opus`, with `cwd: /home/ltms/LTMS/kb`. That matches the operator's rule that a lead is started in the project the fleet works on. **The reviewer separation is weak and the config says so.** With only two usable backends, `gx` reviews diffs that `gx` may have written. That is a real limit of a small host, not an oversight. **Two config comments are now wrong**, and both are the kind of drift a fleet manager would have to prevent: - The `fleet.leaders` block says the `opus` profile is "NOT defined above on purpose". It **is** defined above. - The same block says a lead cannot start until someone completes `claude` `/login` on this host. That login now exists — `~/.claude.json` has an `oauthAccount`. Whether an `opus` lead can actually spawn is a separate question, because #140 says every `subscription: true` claude-code profile fails to spawn. I did **not** re-test #140 on this host. # 5. Credentials — better here than on the Mac `fleet01` holds only four values, in `~/.fleet/secrets.sh` (mode 600): `AI_GATEWAY_TOKEN`, `WORKER_GITEA_TOKEN`, `GITEA_HOST`, `LAVINMQ_URI`. The Mac's store holds dozens, including an AWS AdministratorAccess key. The file is sourced from `~/.zprofile`, which only a **login** shell reads. A herdr member pane on Linux is an interactive but **non-login** shell — measured as a bare `/usr/bin/zsh` in `/proc/<pid>/cmdline`. So members never inherit these values at all. The CB-633 scrub is a second line of defence here rather than the only one. This is the better arrangement and it is worth copying back to the Mac. `fleetd.yaml` contains **no literal secret**. Every credential is referenced by env var name (`tokenEnv`, `gitTokenEnv`, and now `broker.uriEnv`). ## But the git remotes leak the forge token `~/LTMS/kb/.git/config` has the token written into the remote URL: ``` https://<WORKER_GITEA_TOKEN>@git.ltms.dev/akb/kb.git ``` Both worker worktrees under `~/LTMS/.fleet-worktrees/` inherit the same URL. The file is mode 664. This defeats the credential scrub completely. The scrub removes **environment variables**; it cannot remove a token written into a file inside the very repo the member is told to work in. A member that never sees `WORKER_GITEA_TOKEN` in its environment can still read it with `git remote -v`. I found this by accident, and in doing so I printed the token into an operator transcript. **It needs rotating**, and so does `AI_GATEWAY_TOKEN`, which I exposed separately in the same session by using `${VAR:-default}` to test whether it was set — that expands to the value. Both mistakes are the same shape and both are worth designing against: **the credential leaves through a channel nobody classified as a credential channel.** A fleet manager that reports host state will print git remotes, process lists and config files. Every one of those is a leak path. Suggested immediate fix, separate from any manager work: use ssh remotes, or a credential helper, and never `git clone` with the token inline. ## A false alarm that teaches operators to ignore warnings Startup logs this on `fleet01`: ``` startup secret BRIDGED_WORKER_TOKEN: MISSING (profile 'xf' tokenEnv) ``` `xf` has no `tokenEnv` in the config. The name comes from `FleetConfig.Profile`'s compact constructor, which defaults `tokenEnv` to `"BRIDGED_WORKER_TOKEN"` for **every** profile that leaves it unset — including opencode profiles that need no token at all. So the daemon warns, at every boot, about a variable that nobody configured and nothing needs. The report itself is good and I want to keep it. The default is what makes it lie. A profile that needs no token should produce no line. # 6. The shared broker One LavinMQ 2.9.1 container runs **on fleet01**: `fleet-lavinmq`, image `cloudamqp/lavinmq:2.9.1`, volume `fleet-lavinmq-data`, `--restart unless-stopped`, listening on `0.0.0.0:5672` and `0.0.0.0:15672`. | fleet | vhost | user | |---|---|---| | Mac | `/mac` | `fleet-mac` | | fleet01 | `/fleet01` | `fleet-fleet01` | `guest` is `administrator` on `/` only, and was removed from both fleet vhosts. Neither fleet user has a management tag, so both get **401** from the HTTP API. Inspect the broker with `guest:guest` from **on** fleet01. **The broker had to live here, not on the Mac.** Two measured facts forced it: fleet01 cannot reach the Mac, and — until #152 landed today — an unreachable broker at boot stopped the daemon from starting at all. The second is fixed; the first is not. **Isolation is enforced, not just agreed.** I ran the AMQP contract suite twice with the Mac's credential: 8/8 pass against `/mac`, and 8 errors out of 8 against `/fleet01`. The permission boundary is real. Related, and now more likely to matter: #154 — the inbox pulls whole queues into an unbounded in-memory map, so the broker-side limits never apply. # 7. Project state `~/LTMS/kb` is the akb/kb checkout the lead works in. It sits on branch `llm-backend-configurable`, 3 commits behind `origin/main`, with a modified `vendor/graphiti` submodule. Two worker worktrees are left over from earlier runs, both dirty, both on branches that were already merged: - `.fleet-worktrees/57b308-4` → `worker/debrand-code-assumptions-b-57ac39-4` - `.fleet-worktrees/fec975-1` → `worker/debrand-identity-587928-1` Nothing reaps these. `lifecycle.idleTtlSeconds` reaps **members**, not their worktrees. --- # 8. What a fleet manager would have to own This is the point of the ticket. Everything above is one host, set up by hand. Ordered by how much it hurt. **1. Nothing keeps anything alive.** No supervisor for `fleetd`, none for `herdr`. A reboot ends the fleet silently. The Mac solved this with launchd; the Linux answer exists in the repo as `deploy/bridged.service` but is not installed and carries a known secret defect. This is the single biggest gap. **2. A dead lead is never restored, and nothing says so.** `LeadLauncher.ensureLeads()` runs **once**, inline in `Fleetd.main()`. There is no supervision loop. If the lead pane dies an hour later, the slot stays empty until someone restarts the daemon, and no log line reports it. Related: #150 — leads always report `ready: false`, so the status endpoint cannot be used to detect this either. **3. Deployment is manual, per host, and drifts.** `git pull` + `mvn package` + `restart.sh`, typed by a human on each machine. There is no shared artefact and no record of which host runs which commit. Today both hosts match only because I checked. **4. Config drift has no guard.** Different file names per host (`bridged.yaml` vs `fleetd.yaml`), different profile sets, and comments that are already wrong. Config files are gitignored, so a change to what a key *means* can ship green and break a live host. **5. There is no cross-host anything.** Two fleets now share a broker, and that is all they share. `fleet_list` shows only local leads. A Mac lead cannot see, message, or coordinate with a fleet01 lead. Logs are local files on each host. If a fleet manager is meant to answer "what is my whole fleet doing", none of that exists yet. **6. Diagnosis reads the wrong source by default.** Two herdr sessions on one host, and the plain CLI talks to the wrong one. Any status a manager reports must name the socket it read. **7. Credentials leak through channels nobody classified.** Git remote URLs today. Process lists and config dumps tomorrow. A manager that displays host state is a new leak surface by construction, and it should be designed with redaction as a rule rather than as care. **8. Nothing cleans up.** Stale worktrees, stale panes in the other herdr session, 173 orphan queues left on the old Mac-local broker. Each is small. Together they are the reason a host slowly stops being understandable. ## What already works and should not be redesigned - The **login-shell** rule for starting the daemon. It is the difference between working credentials and a silent failure that appears hours later. - Keeping secrets in `~/.zprofile` so **non-login member panes never inherit them**. Defence before the scrub, not instead of it. - **Naming env vars in config instead of values.** With `broker.uriEnv` there is now no literal secret in either host's config. - The **`--check` shape** of `scripts/redeploy-bridged.sh`: read-only, reports whether each named variable resolves, never prints a value. That is the right model for anything a manager runs against a host. ## Not verified - I did not re-test #140 (subscription profiles failing to spawn) on this host, so I cannot say whether an `opus` lead can start now that the login exists. - I did not launch a member and watch a queue appear on `/fleet01`; the broker connection is proven, an end-to-end delegation on this host is not. - I did not measure whether the two leftover worktrees still hold work worth keeping.
Sign in to join this conversation.
1 Participants
Notifications
Due Date
No due date set.
Dependencies

No dependencies set.

Reference: fleet/fleetd#156