redeploy-fleetd.sh is launchd-only: on fleet01 it kills the daemon and leaves TWO running #492

Closed
opened 2026-09-12 03:48:26 +02:00 by ltms · 1 comment
Owner

Found on 2026-09-12 while preparing to test the lead rollover (#480/#489) on fleet01. This blocks bringing that host current, and it silently reproduces an incident that already cost 20 hours here.

The defect

scripts/redeploy-fleetd.sh knows exactly one supervisor: launchd. It has no systemctl, no systemd, and no OS detection of any kind.

$ grep -cE 'systemctl|systemd' scripts/redeploy-fleetd.sh
0

Supervision is decided by one file test:

launchd_installed() { [ -f "$LAUNCHD_PLIST" ]; }

On a Linux host that plist does not exist, so SUPERVISED=0, and the stop branch takes the unsupervised path:

else
  kill "$OLD_PID"
fi

…followed by a manual nohup start of the new jar.

Why that is dangerous, not merely unsupported

fleet01 runs fleetd under a systemd user unit with Restart=on-failure (measured below). The script's own CB-594 comment already establishes the key fact: a SIGTERM'd JVM reports exit 143 even when its shutdown hook runs to completion. Restart=on-failure treats 143 as a failure.

So on fleet01 the sequence is:

  1. the script kills the daemon → exit 143
  2. systemd sees a failure and restarts the old jar
  3. the script then starts its own copy of the new jar

Two daemons on one herdr session. That is the exact shape of the incident recorded here before — two daemons ran for 20 hours and nothing reported it except fleet_list.coordinator.peers[].consumers: 2. Two daemons on one herdr session also kill each other's members.

The script reports ok throughout, because every check it runs (a fresh fleetd listening line, /healthz 200, no new ERROR lines) is satisfied by either daemon.

Measured on fleet01, 2026-09-12

systemctl --user is-active fleetd            -> active
systemctl --user cat fleetd | grep ExecStart -> /bin/zsh -lc "exec java -jar target/fleetd.jar fleetd.yaml"
                                 Restart=    -> on-failure
                        WorkingDirectory=    -> %h/LTMS/fleetd/fleetd
ls ~/Library/LaunchAgents/dev.ltms.fleetd.plist -> No such file or directory
command -v launchctl                          -> MISSING
command -v systemctl shasum sha256sum pgrep curl zsh -> all present

So shasum and pgrep are fine; launchd is the only break.

The fix

Teach the script that "not launchd" and "not supervised" are different answers.

  1. Detect systemd as a first-class supervisor (systemctl --user is-enabled fleetd, or the unit file's presence), and use systemctl --user restart fleetd for the stop/start pair — never kill plus a manual nohup.
  2. When a supervisor is detected but not supported, refuse — do not fall through to the kill path. That fall-through is the whole defect. The safe default for "I cannot tell who supervises this process" is to stop and say so, not to guess.
  3. Keep every existing check. The verify step must still confirm a fresh fleetd listening line and the ERROR scan, and should additionally confirm that exactly one daemon is running afterwards — pgrep -f returning two pids is the symptom this ticket exists for, and nothing currently looks for it.

scripts/test-redeploy-fleetd.sh already sources the script behind its SOURCED guard, so the detection helpers can be tested directly without touching a live daemon.

Related

  • #480, #489, #491 — the rollover work this blocks on fleet01.
  • fleet01 is 91 commits behind its own cached origin/main, and that cache predates #480. LeadRollover is not in its tree at all.
Found on 2026-09-12 while preparing to test the lead rollover (#480/#489) on fleet01. This blocks bringing that host current, and it silently reproduces an incident that already cost 20 hours here. ## The defect `scripts/redeploy-fleetd.sh` knows exactly one supervisor: launchd. It has no `systemctl`, no `systemd`, and no OS detection of any kind. ``` $ grep -cE 'systemctl|systemd' scripts/redeploy-fleetd.sh 0 ``` Supervision is decided by one file test: ```bash launchd_installed() { [ -f "$LAUNCHD_PLIST" ]; } ``` On a Linux host that plist does not exist, so `SUPERVISED=0`, and the stop branch takes the unsupervised path: ```bash else kill "$OLD_PID" fi ``` …followed by a manual `nohup` start of the new jar. ## Why that is dangerous, not merely unsupported fleet01 runs fleetd under a **systemd user unit** with `Restart=on-failure` (measured below). The script's own CB-594 comment already establishes the key fact: a SIGTERM'd JVM reports **exit 143 even when its shutdown hook runs to completion**. `Restart=on-failure` treats 143 as a failure. So on fleet01 the sequence is: 1. the script `kill`s the daemon → exit 143 2. systemd sees a failure and restarts **the old jar** 3. the script then starts **its own** copy of the new jar Two daemons on one herdr session. That is the exact shape of the incident recorded here before — two daemons ran for 20 hours and nothing reported it except `fleet_list.coordinator.peers[].consumers: 2`. Two daemons on one herdr session also kill each other's members. The script reports `ok` throughout, because every check it runs (a fresh `fleetd listening` line, `/healthz` 200, no new ERROR lines) is satisfied by *either* daemon. ## Measured on fleet01, 2026-09-12 ``` systemctl --user is-active fleetd -> active systemctl --user cat fleetd | grep ExecStart -> /bin/zsh -lc "exec java -jar target/fleetd.jar fleetd.yaml" Restart= -> on-failure WorkingDirectory= -> %h/LTMS/fleetd/fleetd ls ~/Library/LaunchAgents/dev.ltms.fleetd.plist -> No such file or directory command -v launchctl -> MISSING command -v systemctl shasum sha256sum pgrep curl zsh -> all present ``` So `shasum` and `pgrep` are fine; launchd is the only break. ## The fix Teach the script that "not launchd" and "not supervised" are different answers. 1. Detect systemd as a first-class supervisor (`systemctl --user is-enabled fleetd`, or the unit file's presence), and use `systemctl --user restart fleetd` for the stop/start pair — never `kill` plus a manual `nohup`. 2. **When a supervisor is detected but not supported, refuse — do not fall through to the kill path.** That fall-through is the whole defect. The safe default for "I cannot tell who supervises this process" is to stop and say so, not to guess. 3. Keep every existing check. The verify step must still confirm a fresh `fleetd listening` line and the ERROR scan, and should additionally confirm that exactly **one** daemon is running afterwards — `pgrep -f` returning two pids is the symptom this ticket exists for, and nothing currently looks for it. `scripts/test-redeploy-fleetd.sh` already sources the script behind its `SOURCED` guard, so the detection helpers can be tested directly without touching a live daemon. ## Related - #480, #489, #491 — the rollover work this blocks on fleet01. - fleet01 is 91 commits behind its own cached `origin/main`, and that cache predates #480. `LeadRollover` is not in its tree at all.
Author
Owner

Fixed and on main at 136312f, via PR #495 (the systemd branch) and PR #499 (the "unclear" state
that closes the gap #495 left). #499 was built on #495's branch, so one merge carried both.

What the script does now. detect_supervisor() returns one of five answers instead of
launchd-or-nothing: launchd, systemd, none, ambiguous (both signals fire), or unclear (a
supervisor looks present but the script cannot tell whether it drives this daemon).
require_drivable_supervisor() refuses on ambiguous, on unclear, and on anything
detect_supervisor did not return — it can no longer fall through to a bare kill. The systemd
branch stops and starts through systemctl --user. assert_single_daemon() runs after the restart
and dies if more than one daemon pid is live, which is the symptom this ticket was filed for and
which none of the existing checks could see.

none now means only what it says: neither supervisor installed, neither loaded, neither probe
errored.

What I measured myself, on the merged script on this Mac:

$ scripts/redeploy-fleetd.sh --check
supervisor detected: launchd

read-only, all three token checks resolve, exit 0. The suite scripts/test-redeploy-fleetd.sh runs
23 test functions, all invoked, exit 0. (The three FAIL: lines in its output are literal printf
from its own internal mutation cells, not suite failures — they confuse every first-time reader,
including me.) bash -n clean.

I also re-ran the detection measurement across all six branches rather than trusting the report:
the unclear branches produce a detail string of 142 and 146 characters, the four others produce 0
with no leakage between them. Harness proof: making detect_supervisor emit only the kind made the
suite exit 1 with a named failure, so the cells can fail.

One thing this Mac cannot verify. There is no systemd here, so the systemd branch is exercised
only through function stubs. The refusal paths are proven; a real systemctl --user stop|start
against a live unit is not. fleet01 is the host that could measure it, and that is their call.

Left open deliberately, each with its own ticket:

  • #493 — the script builds into the live fleetd/target/fleetd.jar path before stopping the daemon.
    Unblocked now that this script work has landed.
  • #504 — four places in the same script where a failed command is still reported as a clean result.

Closing.

Fixed and on `main` at `136312f`, via PR #495 (the systemd branch) and PR #499 (the "unclear" state that closes the gap #495 left). #499 was built on #495's branch, so one merge carried both. **What the script does now.** `detect_supervisor()` returns one of five answers instead of launchd-or-nothing: `launchd`, `systemd`, `none`, `ambiguous` (both signals fire), or `unclear` (a supervisor looks present but the script cannot tell whether it drives this daemon). `require_drivable_supervisor()` refuses on `ambiguous`, on `unclear`, and on anything `detect_supervisor` did not return — it can no longer fall through to a bare `kill`. The systemd branch stops and starts through `systemctl --user`. `assert_single_daemon()` runs after the restart and dies if more than one daemon pid is live, which is the symptom this ticket was filed for and which none of the existing checks could see. `none` now means only what it says: neither supervisor installed, neither loaded, neither probe errored. **What I measured myself, on the merged script on this Mac:** ``` $ scripts/redeploy-fleetd.sh --check supervisor detected: launchd ``` read-only, all three token checks resolve, exit 0. The suite `scripts/test-redeploy-fleetd.sh` runs 23 test functions, all invoked, exit 0. (The three `FAIL:` lines in its output are literal `printf` from its own internal mutation cells, not suite failures — they confuse every first-time reader, including me.) `bash -n` clean. I also re-ran the detection measurement across all six branches rather than trusting the report: the `unclear` branches produce a detail string of 142 and 146 characters, the four others produce 0 with no leakage between them. Harness proof: making `detect_supervisor` emit only the kind made the suite exit 1 with a named failure, so the cells can fail. **One thing this Mac cannot verify.** There is no systemd here, so the systemd branch is exercised only through function stubs. The refusal paths are proven; a real `systemctl --user stop|start` against a live unit is not. fleet01 is the host that could measure it, and that is their call. **Left open deliberately**, each with its own ticket: - #493 — the script builds into the live `fleetd/target/fleetd.jar` path before stopping the daemon. Unblocked now that this script work has landed. - #504 — four places in the same script where a failed command is still reported as a clean result. Closing.
ltms closed this issue 2026-09-12 05:14:39 +02:00
Sign in to join this conversation.
1 Participants
Notifications
Due Date
No due date set.
Dependencies

No dependencies set.

Reference: fleet/fleetd#492