3da44eed63
jar_id() in redeploy-fleetd.sh called shasum directly, which does not exist on GNU coreutils Linux (Debian/Ubuntu/etc.) — there it silently reported an existing jar as "absent" with exit 0, because the missing command made `cut` succeed on empty input and pipefail's failure was then swallowed by the `|| echo "absent"` fallback. The shell test suite hit the same tool at test-redeploy-fleetd.sh:298-299 and died at exit 127 with zero FAIL lines printed — the same shape as a clean pass on the one channel anyone would check. Adds one hash256() helper (prefer sha256sum, fall back to shasum -a 256, same idiom already used in probe-member-credentials.sh) and points jar_id and the test suite's own reference hash at it. jar_id now has three distinct answers instead of two: absent, a hash, or "unhashable" when neither hasher is on PATH — "absent" is never used for a file that exists. Adds a CI job (shell-tests) that runs scripts/test-redeploy-fleetd.sh on ubuntu-latest, gated on the step's own exit code rather than a FAIL-line count, since a suite that dies before running is exactly what a green run also looks like by that count. New tests: test_jar_id_reports_unhashable_when_no_hasher_on_path (stubbed PATH with neither hasher) and test_no_unguarded_macos_only_hasher_calls (a shape check across every script under scripts/, not named lines — #545 already showed this idiom spreading from two sites to six).
1153 lines
67 KiB
Bash
Executable File
1153 lines
67 KiB
Bash
Executable File
#!/usr/bin/env bash
|
|
#
|
|
# Rebuild and restart the fleetd daemon.
|
|
#
|
|
# A merge is not a deployment: the running daemon holds the jar it was started with, so code merged
|
|
# to main does nothing until this runs. See CLAUDE.md -> "Redeploying the daemon".
|
|
#
|
|
# This script exists to turn eight remembered traps into one auditable command:
|
|
#
|
|
# 1. A piped `mvn` hides BUILD FAILURE behind a zero exit, so the build here is never piped.
|
|
# 2. The daemon must start from a LOGIN shell, or the tokens it hands to members are empty:
|
|
# WORKER_GITEA_TOKEN (workers cannot open a PR) and AI_GATEWAY_TOKEN (401 at llm.ltms.dev).
|
|
# Both are read from the DAEMON's own environment at spawn time, so a value added to
|
|
# secrets.sh after startup is absent. Nothing logs this here, so the script checks and says
|
|
# so — and since CB-594, fleetd's own startup log says so too, by env var name.
|
|
# 3. An old daemon that never actually died looks identical from the outside, so the script waits
|
|
# for the process to exit and for the port to free before it starts a new one.
|
|
# 4. "It started" is not "it works": the script polls /healthz until it answers, and reports the
|
|
# herdr protocol number, because healthz can be green while every spawn fails on a protocol
|
|
# mismatch.
|
|
# 5. Restarting under live members drops their tickets, so the script refuses unless you confirm
|
|
# the fleet is drained.
|
|
# 6. CB-594 — the launchd agent (deploy/dev.ltms.fleetd.plist), if installed and loaded, is a
|
|
# SECOND supervisor: its KeepAlive.SuccessfulExit=false restarts the daemon on any nonzero
|
|
# exit, and a bare SIGTERM makes this JVM exit 143 even with its shutdown hook running to
|
|
# completion (measured — see the CB-594 report). A plain `kill` here would race launchd's own
|
|
# restart of the OLD jar. So this script detects whether the agent is loaded and, only then,
|
|
# swaps `kill` + manual `nohup` for `launchctl unload`/`load` — the one supervisor in control
|
|
# at any moment is whichever one you asked to act, never both.
|
|
# 7. fleetd #492 — a systemd --user unit is a THIRD possible supervisor (seen on a second host):
|
|
# Restart=on-failure treats this JVM's SIGTERM exit code (143, per CB-594 above) as a failure
|
|
# too, so a bare `kill` there would race systemd's own restart of the OLD jar exactly like
|
|
# launchd would. This script now tells launchd, systemd, and "genuinely unsupervised" apart as
|
|
# three different answers, drives whichever one it finds through its own control plane
|
|
# (`launchctl` / `systemctl --user`), and REFUSES outright — never falls back to `kill` — when
|
|
# it finds a supervision signal it cannot map to exactly one of the two it knows how to drive.
|
|
# A wrong guess here is how two daemons end up running against one herdr session. Follow-up:
|
|
# "not currently loaded" is not the same fact as "unsupervised" — a unit that is installed but
|
|
# activating/failed/pending-restart, or a `systemctl` call that could not answer at all (e.g.
|
|
# no user-bus access), both now read as a fifth answer, "unclear", and REFUSE the same way
|
|
# "ambiguous" does, rather than silently falling through to "none".
|
|
# 8. fleetd #492 — a post-restart check counts running fleetd processes and fails the whole run if
|
|
# more than one is alive. That is the one thing none of the checks above (healthz 200, jar id,
|
|
# the fresh "listening" line) can see: every one of them is satisfied by EITHER daemon.
|
|
# 9. fleetd #512 — the ERROR-line count above is blind by construction to the exact failure #493
|
|
# is about: an uncaught exception in a shutdown thread never passes through the logger, so it
|
|
# never carries an ERROR (or SEVERE) token that any count could see. This script now also
|
|
# greps the previous daemon's shutdown window for that exception's real shape, and separately
|
|
# asserts that SessionManager's drain-complete line (fleetd #522) is present there — its
|
|
# absence is the real signal, because a drain that dies on its first session prints nothing
|
|
# else either. Warns loudly; never fails the redeploy, because by the time this is detectable
|
|
# the new daemon is already up and healthy.
|
|
#
|
|
# Usage:
|
|
# scripts/redeploy-fleetd.sh # build, confirm, restart, verify
|
|
# scripts/redeploy-fleetd.sh --yes # skip the drain confirmation (fleet already checked)
|
|
# scripts/redeploy-fleetd.sh --no-build # restart the jar already on disk
|
|
# scripts/redeploy-fleetd.sh --check # report state and exit; changes nothing
|
|
#
|
|
# Exits non-zero on any failure. A failed build never stops the running daemon.
|
|
|
|
set -euo pipefail
|
|
|
|
REPO="$(cd "$(dirname "${BASH_SOURCE[0]}")/.." && pwd)"
|
|
MODULE="$REPO/fleetd"
|
|
JAR="$MODULE/target/fleetd.jar"
|
|
# fleetd #493: never build into the path a running process holds. The build writes here first
|
|
# (Maven's shade plugin has finalName=fleetd, so `clean install` still lands its output at
|
|
# target/fleetd.jar — that part is unchanged and out of this script's control), but this script
|
|
# now moves it out to JAR_STAGED immediately, and only swaps it back to JAR (a plain `mv`, so a
|
|
# rename, never a byte-by-byte overwrite) after the OLD daemon has been confirmed exited. See
|
|
# stage_built_jar/swap_staged_jar below.
|
|
JAR_STAGED="$MODULE/target/fleetd-new.jar"
|
|
OUT="$MODULE/fleetd.out"
|
|
# Matches BOTH the absolute form and the relative `java -jar target/fleetd.jar` a hand-start
|
|
# produces from inside fleetd/. Anchoring on the absolute path alone was a real bug: the daemon
|
|
# restarted correctly and the script still reported "no process appeared", because it launched with
|
|
# a relative path and then looked for an absolute one.
|
|
PATTERN='target/fleetd.jar'
|
|
HEALTH='http://127.0.0.1:8765/healthz'
|
|
STOP_WAIT=30 # seconds to wait for a clean exit before reporting failure
|
|
HEALTH_WAIT=60 # seconds to wait for /healthz to answer after start
|
|
|
|
# CB-594: the launchd agent this script must not fight with (see trap 6 above).
|
|
LAUNCHD_LABEL='dev.ltms.fleetd'
|
|
LAUNCHD_PLIST="$HOME/Library/LaunchAgents/$LAUNCHD_LABEL.plist"
|
|
|
|
# fleetd #492: the systemd --user unit this script must not fight with either (see trap 7 above).
|
|
# Measured on the second host: `systemctl --user cat fleetd` names the unit "fleetd" (not
|
|
# "dev.ltms.fleetd" — systemd user units here are not namespaced the way the launchd label is).
|
|
SYSTEMD_UNIT='fleetd'
|
|
|
|
# fleetd #492 follow-up: detect_supervisor packs TWO values (kind, detail) onto the one stdout
|
|
# line that survives its $(...) call — see the constraints comment above that function. This is
|
|
# the separator between them: the ASCII "unit separator" byte, chosen because it never occurs in
|
|
# any of the prose detail strings and needs no escaping in a `case`/glob pattern.
|
|
SUPERVISOR_DETAIL_SEP=$'\x1f'
|
|
|
|
# fleetd #492 follow-up, refined by fleetd #545: three states, not two, set by
|
|
# systemd_loaded/systemd_installed —
|
|
# 0 = no error, the probe ran and gave a clean answer.
|
|
# 1 = the probe RAN and answered badly: `systemctl` exited non-zero AND wrote something to
|
|
# stderr, a real tool failure (e.g. it cannot reach the user bus over a non-lingering ssh
|
|
# session), never the same fact as a clean negative answer ("not active", no stderr).
|
|
# 2 = the probe could not even be SET UP: the `mktemp` call that makes a place to capture
|
|
# `systemctl`'s stderr failed before `systemctl` ever ran. This is a fleetd #545 fix: on GNU
|
|
# coreutils (every Linux distribution) a template with no `X`s made `mktemp` fail every
|
|
# single time, and the two states were folded into one flag and one message that named
|
|
# cause 1 ("systemctl exited non-zero and reported an error on stderr") for a failure that
|
|
# was actually cause 2 — systemctl was never executed at all. One flag with two meanings
|
|
# needing different messages was the defect; a third value is the fix, not a second flag.
|
|
# Initialized here, not just inside the probes, so detect_supervisor can read them under `set -u`
|
|
# even before either probe has ever run, and so a test that stubs a probe with a plain
|
|
# `return 0`/`return 1` body (leaving these untouched) reads a deterministic 0 rather than whatever
|
|
# a previous probe call left behind.
|
|
SYSTEMD_LOADED_ERRORED=0
|
|
SYSTEMD_INSTALLED_ERRORED=0
|
|
# fleetd #492 follow-up: SUPERVISOR_UNCLEAR_DETAIL is the specific supervisor/reason that
|
|
# require_drivable_supervisor's die() names on an "unclear" answer. Deliberately NOT pre-declared
|
|
# here (unlike the two flags above): it is set only by the real call site, right after it unpacks
|
|
# detect_supervisor's stdout (see the constraints comment above detect_supervisor). If that call
|
|
# site is ever skipped or broken, a bare `set -u` reference to this variable in
|
|
# require_drivable_supervisor must fail loudly with "unbound variable" — a pre-declared empty
|
|
# default would instead silently print an empty reason, hiding exactly the value this ticket
|
|
# exists to surface.
|
|
|
|
DO_BUILD=1; ASSUME_YES=0; CHECK_ONLY=0
|
|
for arg in "$@"; do
|
|
case "$arg" in
|
|
--yes|-y) ASSUME_YES=1 ;;
|
|
--no-build) DO_BUILD=0 ;;
|
|
--check) CHECK_ONLY=1 ;;
|
|
-h|--help) sed -n '3,48p' "${BASH_SOURCE[0]}"; exit 0 ;;
|
|
*) echo "unknown option: $arg (try --help)" >&2; exit 2 ;;
|
|
esac
|
|
done
|
|
|
|
say() { printf '\n\033[1m== %s\033[0m\n' "$*"; }
|
|
ok() { printf ' ok %s\n' "$*"; }
|
|
warn() { printf ' WARN %s\n' "$*"; }
|
|
die() { printf '\n FAIL %s\n\n' "$*" >&2; exit 1; }
|
|
|
|
# fleetd #550 — shasum is macOS-only (it ships with Perl, which Debian/Ubuntu/etc. do not install
|
|
# by default); GNU coreutils (every mainstream Linux distro) ships sha256sum instead and has no
|
|
# shasum at all. Prefer sha256sum, fall back to shasum -a 256 — same idiom as
|
|
# probe-member-credentials.sh's `hasher` selection — and when NEITHER is on PATH, say so plainly.
|
|
# That third answer matters: without it, a missing hasher makes `cut` succeed on empty input, and
|
|
# under `set -o pipefail` the pipeline as a whole still fails, so a caller's own `|| echo "absent"`
|
|
# then reports a file that is right there as though it were gone. `hash256` never does that — it
|
|
# only ever hashes or says it could not.
|
|
hash256() {
|
|
local f="$1"
|
|
if command -v sha256sum >/dev/null 2>&1; then
|
|
sha256sum "$f" | cut -c1-12
|
|
elif command -v shasum >/dev/null 2>&1; then
|
|
shasum -a 256 "$f" | cut -c1-12
|
|
else
|
|
echo "unhashable"
|
|
fi
|
|
}
|
|
|
|
# Reports the hash of $JAR by default, or of whatever path is passed — used to report the STAGED
|
|
# jar right after a build (before it has been swapped in) without ever changing what a bare
|
|
# `jar_id` (no args) means: the live path, $JAR. --check and the final "pid ..., jar ..." line
|
|
# both call it with no args on purpose, so neither can ever be fooled by a leftover staged file.
|
|
# fleetd #550 — THREE distinct answers now, not two: `[ -f "$f" ]` already separates "the jar is
|
|
# not there" (-> "absent") from "the jar is there"; for the second case, hash256 itself separates
|
|
# "hashed it" (a 12-char hex string) from "could not hash it" (-> "unhashable", when no hasher is
|
|
# on PATH). "absent" must never be the answer for a file that exists — that conflation, on Linux,
|
|
# was the whole defect this ticket fixes.
|
|
jar_id() { local f="${1:-$JAR}"; [ -f "$f" ] && hash256 "$f" || echo "absent"; }
|
|
running_pid() { pgrep -f "$PATTERN" || true; }
|
|
|
|
# fleetd #493 — three small, independently testable pieces of "never build into the path a
|
|
# running process holds":
|
|
#
|
|
# stage_built_jar moves the jar Maven just produced OUT of the live path and onto the staging
|
|
# path, immediately after a successful build. Dies (leaving the OLD daemon
|
|
# untouched — this runs before the stop step) if Maven reported success but
|
|
# left no jar behind, or if the move itself fails.
|
|
# require_no_build_jar the --no-build path never builds or stages anything: it must find a
|
|
# jar already sitting at the live path from an earlier successful run, and
|
|
# die with the same truthful message this script has always used if not.
|
|
# wait_for_daemon_exit polls running_pid() for up to $1 seconds and reports whether the OLD
|
|
# daemon actually exited — extracted to its own function so the main flow
|
|
# can be relied on to call swap_staged_jar only AFTER this returns success,
|
|
# and so a test can prove that ordering by reading the script's own source.
|
|
# swap_staged_jar the actual swap: a plain `mv` of the staged jar onto the live path. Called
|
|
# only once the OLD daemon is confirmed gone (see wait_for_daemon_exit above),
|
|
# so this is never a write into a path a running process holds — by the time
|
|
# it runs, nothing holds that path anymore. If it fails, the caller must not
|
|
# start a new daemon: die() below already refuses that by exiting the script.
|
|
stage_built_jar() {
|
|
[ -f "$JAR" ] || die "build succeeded but produced no jar at $JAR — cannot stage it for restart.
|
|
The running daemon was NOT touched."
|
|
mv -f "$JAR" "$JAR_STAGED" \
|
|
|| die "could not move the freshly built jar from $JAR to the staging path $JAR_STAGED.
|
|
The running daemon was NOT touched."
|
|
}
|
|
|
|
require_no_build_jar() {
|
|
[ -f "$JAR" ] || die "no jar at $JAR — run without --no-build"
|
|
}
|
|
|
|
wait_for_daemon_exit() {
|
|
local timeout="$1" _i
|
|
for _i in $(seq "$timeout"); do
|
|
[ -z "$(running_pid)" ] && return 0
|
|
sleep 1
|
|
done
|
|
[ -z "$(running_pid)" ]
|
|
}
|
|
|
|
swap_staged_jar() {
|
|
local staged="$1" live="$2"
|
|
[ -f "$staged" ] || die "no staged jar at $staged to swap in — the daemon was NOT started."
|
|
mv -f "$staged" "$live" \
|
|
|| die "could not move the staged jar from $staged into place at $live — the daemon was NOT
|
|
started. The built jar is still sitting at $staged; a manual 'mv \"$staged\" \"$live\"'
|
|
may recover this once you find out why the move failed."
|
|
}
|
|
|
|
# fleetd #521 — the swap decision, and the step that acts on it.
|
|
#
|
|
# The defect: the swap step used to be guarded inline by `if [ "$DO_BUILD" = 1 ]` in the main flow.
|
|
# Changing that to `if false` left the suite green and the swap never ran, so a redeploy reported
|
|
# every step succeeding while the daemon started on no jar at all (stage_built_jar has already moved
|
|
# the freshly built one to $JAR_STAGED by then) or on a stale one.
|
|
# test_swap_ordered_after_wait_and_before_start could not catch it: it reads this script's own text
|
|
# and compares line positions, and a same-line edit moves no line.
|
|
#
|
|
# Why these are TWO functions, and why the second one exists at all. Extracting only the predicate
|
|
# — `should_swap`, which is what #521 asked for — is not enough, and this was measured, not guessed:
|
|
# with the main flow calling `if should_swap "$DO_BUILD"; then`, changing THAT to `if false; then`
|
|
# still left the whole suite at exit 0 with no failures. Tests that call a predicate directly prove
|
|
# the predicate is right; nothing makes the code that does the work consult it. Extraction had moved
|
|
# the untested decision one level up rather than removing it.
|
|
#
|
|
# So the decision and the action live together in swap_if_built, and the main flow has no guard of
|
|
# its own to get wrong — it calls one function unconditionally. A test then calls swap_if_built with
|
|
# both values of do_build and checks whether the swap actually happened, which fails if the guard is
|
|
# removed, inverted, or stops being consulted. should_swap stays a separate predicate because it is
|
|
# the decision itself and is worth naming and testing on its own.
|
|
#
|
|
# What this still does not pin: deleting the swap_if_built call from the main flow altogether. That
|
|
# is the ordering test's job — its needle is that call site — and no test in this file can do better,
|
|
# because sourcing stops before the main flow ever runs (see the SOURCED guard below).
|
|
should_swap() {
|
|
local do_build="$1"
|
|
[ "$do_build" = 1 ]
|
|
}
|
|
|
|
swap_if_built() {
|
|
local do_build="$1"
|
|
should_swap "$do_build" || return 0
|
|
say "swap"
|
|
swap_staged_jar "$JAR_STAGED" "$JAR"
|
|
ok "jar in place: $(jar_id)"
|
|
}
|
|
|
|
# `launchctl list <label>` exits 0 iff the label is loaded (registered with launchd) — true whether
|
|
# or not it is currently running, which is exactly "supervision is active" for our purposes. Read-
|
|
# only: neither helper below changes anything, so both are also safe under --check.
|
|
launchd_installed() { [ -f "$LAUNCHD_PLIST" ]; }
|
|
launchd_loaded() { launchctl list "$LAUNCHD_LABEL" >/dev/null 2>&1; }
|
|
|
|
# fleetd #492: same two questions for systemd --user. Kept as separate, overridable functions
|
|
# (never an inline `systemctl` call at each use site) so a test on a box with no systemd at all
|
|
# (this repo is developed on macOS) can substitute each one independently — the same seam
|
|
# launchd_installed/launchd_loaded above already use.
|
|
#
|
|
# fleetd #492 follow-up: both functions used to throw `systemctl`'s stderr straight into
|
|
# /dev/null, which meant "systemctl answered no" and "systemctl could not answer at all" (e.g. it
|
|
# cannot reach the user bus over a non-lingering ssh session) looked identical — both a plain
|
|
# nonzero exit. They now capture stderr separately and set their own *_ERRORED flag to 1 when the
|
|
# call exited non-zero AND wrote something to stderr — a real tool failure, never a clean "not
|
|
# installed"/"not active" answer (which exits non-zero with empty stderr). detect_supervisor reads
|
|
# the flag right after calling the probe, so a probe that could not answer routes to "unclear",
|
|
# never silently becomes "none".
|
|
#
|
|
# fleetd #545: the flag has a third value, 2, set when the `mktemp` call that sets up the probe's
|
|
# own stderr capture fails, before `systemctl` ever runs — see the SYSTEMD_LOADED_ERRORED /
|
|
# SYSTEMD_INSTALLED_ERRORED comment above their initialization for why this is a third value on the
|
|
# same flag, not a second flag.
|
|
#
|
|
# "installed": a unit FILE by this name exists, regardless of its current state — the systemd
|
|
# analogue of the plist file existing on disk. `list-unit-files` reads unit definitions without
|
|
# depending on runtime state, so this stays read-only and safe under --check.
|
|
systemd_installed() {
|
|
SYSTEMD_INSTALLED_ERRORED=0
|
|
command -v systemctl >/dev/null 2>&1 || return 1
|
|
local err_file out rc=0
|
|
if ! err_file="$(mktemp -t systemd-installed-err.XXXXXX)"; then
|
|
SYSTEMD_INSTALLED_ERRORED=2
|
|
return 1
|
|
fi
|
|
out="$(systemctl --user list-unit-files "$SYSTEMD_UNIT.service" --no-legend 2>"$err_file")" || rc=$?
|
|
if [ "$rc" -ne 0 ]; then
|
|
if [ -s "$err_file" ]; then
|
|
SYSTEMD_INSTALLED_ERRORED=1
|
|
fi
|
|
rm -f "$err_file"
|
|
return "$rc"
|
|
fi
|
|
rm -f "$err_file"
|
|
printf '%s' "$out" | grep -q .
|
|
}
|
|
# "loaded": systemd currently supervises this unit as an active job — the systemd analogue of
|
|
# `launchctl list <label>` succeeding. Measured on the second host: `systemctl --user is-active
|
|
# fleetd` -> "active". A clean "no" (inactive/failed/activating/deactivating) exits non-zero with
|
|
# nothing on stderr; a probe that could not reach systemd at all exits non-zero WITH a stderr
|
|
# message — see the fleetd #492 follow-up note above.
|
|
systemd_loaded() {
|
|
SYSTEMD_LOADED_ERRORED=0
|
|
command -v systemctl >/dev/null 2>&1 || return 1
|
|
local err_file rc=0
|
|
if ! err_file="$(mktemp -t systemd-loaded-err.XXXXXX)"; then
|
|
SYSTEMD_LOADED_ERRORED=2
|
|
return 1
|
|
fi
|
|
systemctl --user is-active "$SYSTEMD_UNIT" >/dev/null 2>"$err_file" || rc=$?
|
|
if [ "$rc" -ne 0 ] && [ -s "$err_file" ]; then
|
|
SYSTEMD_LOADED_ERRORED=1
|
|
fi
|
|
rm -f "$err_file"
|
|
return "$rc"
|
|
}
|
|
|
|
# fleetd #504: the "loaded but not currently running" branches in the main stop step (case
|
|
# launchd/systemd, reached when $OLD_PID is empty) used to run `launchctl unload`/`systemctl --user
|
|
# stop` with `2>/dev/null || true` and then print `ok` unconditionally — the exact conflation
|
|
# systemd_loaded/systemd_installed above already fixed on the READ side (fleetd #492 follow-up): a
|
|
# genuine "already stopped" answer (nonzero exit, nothing on stderr) is harmless, but a real tool
|
|
# failure (nonzero exit WITH a stderr message — e.g. launchd or the systemd user bus is
|
|
# unreachable) is not, and reporting `ok` on THAT means the start step below can register a fresh
|
|
# load on top of a supervisor that never actually let go: the exact two-daemons failure fleetd #492
|
|
# exists to prevent, reached from the one state (already odd) where a false `ok` is least
|
|
# affordable. These two functions apply the same "capture stderr separately, flag only a nonzero
|
|
# exit WITH stderr as a real failure" pattern to the WRITE side. No ${VAR:-default} anywhere here —
|
|
# see the #492 follow-up constraints comment above detect_supervisor for why a default would hide a
|
|
# lost value instead of surfacing it (fleetd #497's defect class).
|
|
unload_launchd_if_loaded() {
|
|
local err_file rc=0
|
|
if ! err_file="$(mktemp -t launchd-unload-err.XXXXXX)"; then
|
|
die "could not create a temp file to capture 'launchctl unload' stderr — cannot tell a real
|
|
failure from a clean already-unloaded answer, so refusing to guess. The daemon's
|
|
supervision state was NOT touched."
|
|
fi
|
|
launchctl unload -w "$LAUNCHD_PLIST" >/dev/null 2>"$err_file" || rc=$?
|
|
if [ "$rc" -ne 0 ] && [ -s "$err_file" ]; then
|
|
die "'launchctl unload -w $LAUNCHD_PLIST' failed: $(cat "$err_file")
|
|
The daemon may still be under supervision; investigate before retrying."
|
|
fi
|
|
rm -f "$err_file"
|
|
}
|
|
|
|
stop_systemd_if_loaded() {
|
|
local err_file rc=0
|
|
if ! err_file="$(mktemp -t systemd-stop-err.XXXXXX)"; then
|
|
die "could not create a temp file to capture 'systemctl --user stop' stderr — cannot tell a
|
|
real failure from a clean already-stopped answer, so refusing to guess. The daemon's
|
|
supervision state was NOT touched."
|
|
fi
|
|
systemctl --user stop "$SYSTEMD_UNIT" >/dev/null 2>"$err_file" || rc=$?
|
|
if [ "$rc" -ne 0 ] && [ -s "$err_file" ]; then
|
|
die "'systemctl --user stop $SYSTEMD_UNIT' failed: $(cat "$err_file")
|
|
The daemon may still be under supervision; investigate before retrying."
|
|
fi
|
|
rm -f "$err_file"
|
|
}
|
|
|
|
# fleetd #492: three real answers, not two — launchd, systemd, or genuinely unsupervised — plus a
|
|
# fourth, "ambiguous", for the one case this script cannot tell apart: both signals firing at once.
|
|
# That is exactly "I cannot tell who supervises this process", and guessing wrong here is how two
|
|
# daemons end up running against one herdr session (see trap 7 in the header).
|
|
#
|
|
# fleetd #492 follow-up: a fifth answer, "unclear", for two more situations that must NEVER be read
|
|
# as "none" (measured — see the report this ticket is a follow-up to):
|
|
# - installed-but-not-loaded, on EITHER supervisor. `systemctl --user is-active` answers "no" for
|
|
# `activating`, `deactivating`, `failed`, and while an auto-restart is pending — every one of
|
|
# those is a host that IS under systemd (or launchd) and whose supervisor is about to act again.
|
|
# `*_installed` already knows the unit/agent exists; this is the first place that fact is
|
|
# actually consulted in the decision, not just printed as a warning.
|
|
# - a probe that could not answer at all. systemd_loaded/systemd_installed set their own
|
|
# *_ERRORED flag (see the comment above them) when `systemctl` exits non-zero WITH a stderr
|
|
# message — a real tool failure, e.g. it cannot reach the user bus over a non-lingering ssh
|
|
# session — never conflated with a clean negative answer.
|
|
# "none" now means only: neither supervisor is installed, neither is loaded, and neither probe
|
|
# errored.
|
|
#
|
|
# fleetd #545: *_ERRORED carries a THIRD state (2 = the probe's own mktemp setup failed, before
|
|
# `systemctl` ever ran — see the flag's own comment above its initialization), and it must never be
|
|
# reported with the same detail text as state 1 (`systemctl` ran and answered badly on stderr). The
|
|
# two are different facts about different failures, and conflating them makes the "unclear" message
|
|
# assert a cause ("systemctl exited non-zero and reported an error on stderr") that was never
|
|
# measured when the real cause was state 2. detect_supervisor below picks the detail text off the
|
|
# flag's value, not off a single "errored at all" boolean.
|
|
#
|
|
# fleetd #492 follow-up — constraints every caller of this function depends on (learned the hard
|
|
# way: an earlier version of this fix set a SUPERVISOR_UNCLEAR_DETAIL global from inside here and
|
|
# it was silently lost, because every real call site invokes this as `$(detect_supervisor)`):
|
|
# 1. It is called as `$(detect_supervisor)`, so ONLY STDOUT crosses back to the caller. Anything
|
|
# this function needs to tell its caller — the "unclear" detail included — must be printed,
|
|
# never assigned to a global: a global set inside a `$( )` subshell dies with that subshell.
|
|
# This function packs BOTH values (kind and detail) onto that one stdout line, joined by
|
|
# $SUPERVISOR_DETAIL_SEP, and the caller unpacks them on its own side of the subshell boundary.
|
|
# 2. This script runs under `set -u` (part of the `set -euo pipefail` at the top of the file), so
|
|
# an unset variable is a loud failure. Do not add a `${VAR:-default}` anywhere downstream to
|
|
# paper over a value that should always be there — that hides a lost value instead of
|
|
# surfacing it (fleetd #497's defect class). The mechanism is named rather than cited by line
|
|
# number on purpose: a line number in a comment goes stale on the next insert above it, and
|
|
# this one already had — it said line 50 while the `set` line was at 54.
|
|
# 3. Every `case` on this function's return value needs an explicit final `*)` arm, chosen by
|
|
# whether that caller ACTS on the value (`die` — an unrecognised value must never be silently
|
|
# driven) or only DISPLAYS it (`echo`/`warn` and continue — a diagnostic must not go silent on
|
|
# exactly the value it most needs to report).
|
|
#
|
|
# Pure and side-effect-free besides the two *_ERRORED flags (read back within this same call, never
|
|
# by the caller — see the constraints above): reads the four probes and decides — never mutates
|
|
# anything, so it is safe under --check and testable by overriding
|
|
# launchd_installed/launchd_loaded/systemd_installed/systemd_loaded after sourcing.
|
|
detect_supervisor() {
|
|
local ld=0 sd=0 li=0 si=0 kind detail=""
|
|
SYSTEMD_LOADED_ERRORED=0
|
|
SYSTEMD_INSTALLED_ERRORED=0
|
|
|
|
launchd_loaded && ld=1
|
|
systemd_loaded && sd=1
|
|
launchd_installed && li=1
|
|
systemd_installed && si=1
|
|
|
|
if [ "$SYSTEMD_LOADED_ERRORED" = 2 ] || [ "$SYSTEMD_INSTALLED_ERRORED" = 2 ]; then
|
|
detail="the systemd --user probe for '$SYSTEMD_UNIT' could not even be set up (a temp file to capture systemctl's stderr could not be created) — systemctl was never run, so this says nothing about systemd, the user bus, or the unit itself"
|
|
kind="unclear"
|
|
elif [ "$SYSTEMD_LOADED_ERRORED" = 1 ] || [ "$SYSTEMD_INSTALLED_ERRORED" = 1 ]; then
|
|
detail="the systemd --user probe for '$SYSTEMD_UNIT' could not answer cleanly (systemctl exited non-zero and reported an error on stderr, not a clean negative — e.g. it cannot reach the user bus)"
|
|
kind="unclear"
|
|
elif [ "$ld" = 1 ] && [ "$sd" = 1 ]; then
|
|
kind="ambiguous"
|
|
elif [ "$li" = 1 ] && [ "$ld" = 0 ]; then
|
|
detail="the launchd agent ($LAUNCHD_LABEL) is installed ($LAUNCHD_PLIST exists) but is not currently loaded"
|
|
kind="unclear"
|
|
elif [ "$si" = 1 ] && [ "$sd" = 0 ]; then
|
|
detail="the systemd --user unit ($SYSTEMD_UNIT) is installed but not currently active — it may be activating, deactivating, failed, or waiting on an auto-restart"
|
|
kind="unclear"
|
|
elif [ "$ld" = 1 ]; then
|
|
kind="launchd"
|
|
elif [ "$sd" = 1 ]; then
|
|
kind="systemd"
|
|
else
|
|
kind="none"
|
|
fi
|
|
|
|
printf '%s%s%s' "$kind" "$SUPERVISOR_DETAIL_SEP" "$detail"
|
|
}
|
|
|
|
# fleetd #492: turns anything detect_supervisor returns that is NOT exactly one of the two
|
|
# supervisors this script knows how to drive into a die() — never a fall-through to the `kill`
|
|
# path. Kept as its own function so a test can call it directly (in a subshell, since it die()s)
|
|
# without running the whole report-state flow or needing a real launchd/systemd.
|
|
require_drivable_supervisor() {
|
|
local kind="$1"
|
|
case "$kind" in
|
|
launchd|systemd|none) ;;
|
|
ambiguous)
|
|
die "both launchd ($LAUNCHD_LABEL) and systemd --user ($SYSTEMD_UNIT) report themselves as
|
|
loaded for this daemon at the same time. This script cannot tell which one actually
|
|
supervises the running process, and driving either alone risks the OTHER reviving the
|
|
OLD jar out from under it — the exact failure this ticket (fleetd #492) exists to
|
|
prevent. Stop one of the two supervisors by hand, confirm only one remains loaded, then
|
|
rerun." ;;
|
|
unclear)
|
|
# fleetd #492 follow-up: SUPERVISOR_UNCLEAR_DETAIL crosses back from detect_supervisor's
|
|
# subshell via its stdout, unpacked by the caller BEFORE it calls this function (see the
|
|
# constraints comment above detect_supervisor). No ${VAR:-default} here on purpose: if the
|
|
# detail is somehow missing, `set -u` makes this reference fail loudly instead of silently
|
|
# naming nothing — a default that hides a lost value is the same defect class as fleetd
|
|
# #497.
|
|
die "a supervisor looks present but this script cannot tell whether it actually drives this
|
|
daemon: $SUPERVISOR_UNCLEAR_DETAIL. Guessing wrong here is the same failure 'ambiguous'
|
|
above exists to prevent: driving the daemon while an unseen supervisor revives the OLD
|
|
jar out from under it (fleetd #492). Check 'launchctl list $LAUNCHD_LABEL' and
|
|
'systemctl --user status $SYSTEMD_UNIT' by hand, resolve whichever looks unclear, then
|
|
rerun." ;;
|
|
*)
|
|
die "detect_supervisor returned an unrecognized value '$kind' — refusing to guess which
|
|
supervisor, if any, controls this daemon." ;;
|
|
esac
|
|
}
|
|
|
|
# fleetd #492: the exact symptom a racing supervisor produces — count how many fleetd processes are
|
|
# alive right now. Takes the pid list as a parameter (rather than calling running_pid() itself) so a
|
|
# test can pass a canned two-line string without a real second process running. Pure except for the
|
|
# die() in assert_single_daemon below.
|
|
count_daemon_pids() {
|
|
local pids="$1"
|
|
if [ -z "$pids" ]; then
|
|
echo 0
|
|
else
|
|
printf '%s\n' "$pids" | grep -c .
|
|
fi
|
|
}
|
|
assert_single_daemon() {
|
|
local pids="$1" count
|
|
count="$(count_daemon_pids "$pids")"
|
|
if [ "$count" -gt 1 ]; then
|
|
die "more than one fleetd process is running after this restart (pids: $(printf '%s' "$pids" | tr '\n' ' ')).
|
|
This is the exact failure a racing supervisor produces: the OLD jar was revived by its
|
|
supervisor while this script started a NEW copy. Two daemons on one herdr session kill
|
|
each other's members. Investigate with 'pgrep -f \"$PATTERN\"' and stop the wrong one by
|
|
hand — do not assume either pid is the one you want."
|
|
fi
|
|
}
|
|
|
|
# CB-600: the script computes its own log path from where it sits on disk (REPO, above); the
|
|
# plist hard-codes an absolute StandardOutPath. Nothing forced the two to agree — if this script
|
|
# were ever run from a checkout other than the one the loaded plist names, launchd would start and
|
|
# log the daemon correctly, while every check below (the fresh "fleetd listening" line, the
|
|
# ERROR-count scan) would read a different, empty or stale file and the script would report a
|
|
# clean restart while the daemon crash-loops. Pure and side-effect-free besides `die`/`ok` — reads
|
|
# the two paths, resolves them, compares — so it never touches launchd or the daemon and can be
|
|
# exercised by sourcing this script (see the SOURCED guard below) without installing the agent.
|
|
check_log_path_matches_plist() {
|
|
local script_out="$1" plist_path="$2"
|
|
local plist_out resolved_out resolved_plist_out
|
|
# Checked by exit status, not by emptiness: on a missing file/key PlistBuddy exits nonzero but
|
|
# still writes a message ("File Doesn't Exist, Will Create: ...") that command substitution
|
|
# would happily capture as if it were the real value — testing only `-z` missed that case.
|
|
if ! plist_out="$(/usr/libexec/PlistBuddy -c 'Print :StandardOutPath' "$plist_path" 2>/dev/null)" \
|
|
|| [ -z "$plist_out" ]; then
|
|
die "launchd agent is loaded but PlistBuddy could not read StandardOutPath from
|
|
$plist_path
|
|
— cannot verify the daemon logs where this script is about to look. Fix the plist before
|
|
redeploying supervised."
|
|
fi
|
|
resolved_out="$(cd "$(dirname "$script_out")" 2>/dev/null && pwd -P)/$(basename "$script_out")" || true
|
|
resolved_plist_out="$(cd "$(dirname "$plist_out")" 2>/dev/null && pwd -P)/$(basename "$plist_out")" || true
|
|
if [ -z "$resolved_out" ] || [ -z "$resolved_plist_out" ] || [ "$resolved_out" != "$resolved_plist_out" ]; then
|
|
die "log path mismatch — this script reads
|
|
$script_out (resolved: ${resolved_out:-<directory does not exist>})
|
|
but the loaded plist's StandardOutPath is
|
|
$plist_out (resolved: ${resolved_plist_out:-<directory does not exist>})
|
|
Under supervision the daemon writes to the PLIST's path, not necessarily this script's — every
|
|
post-restart check below (the fresh 'fleetd listening' line, the ERROR-count scan) would read
|
|
the wrong file and could report a clean restart while the daemon crash-loops. Fix the mismatch
|
|
(move this checkout to match the plist, or edit the plist's StandardOutPath/StandardErrorPath)
|
|
before redeploying supervised."
|
|
fi
|
|
ok "log path check: script and plist agree ($resolved_out)"
|
|
}
|
|
|
|
# Classify ERROR lines in one fresh log region. AMQP failure messages now include the connection
|
|
# name, so a recovery can clear only errors for its own connection. A candidate with neither name
|
|
# remains unexplained: it must never be quieted by a recovery on the other connection.
|
|
classify_amqp_connection_errors() {
|
|
local log_file="$1" line pending_inbox=0 pending_lead_mailbox=0
|
|
REDEPLOY_ERROR_COUNT=0
|
|
REDEPLOY_RECOVERED_AMQP_ERRORS=0
|
|
REDEPLOY_UNEXPLAINED_ERRORS=0
|
|
|
|
while IFS= read -r line || [ -n "$line" ]; do
|
|
case "$line" in
|
|
*' ERROR '*|*' SEVERE '*)
|
|
REDEPLOY_ERROR_COUNT=$((REDEPLOY_ERROR_COUNT + 1))
|
|
case "$line" in
|
|
*'AMQP connection fleetd-reply-inbox: An unexpected connection driver error occurred'*|*'AMQP connection fleetd-reply-inbox: Caught an exception during connection recovery!'*)
|
|
pending_inbox=$((pending_inbox + 1))
|
|
;;
|
|
*'AMQP connection fleetd-lead-mailbox: An unexpected connection driver error occurred'*|*'AMQP connection fleetd-lead-mailbox: Caught an exception during connection recovery!'*)
|
|
pending_lead_mailbox=$((pending_lead_mailbox + 1))
|
|
;;
|
|
*'AMQP connection'*'An unexpected connection driver error occurred'*|*'AMQP connection'*'Caught an exception during connection recovery!'*)
|
|
REDEPLOY_UNEXPLAINED_ERRORS=$((REDEPLOY_UNEXPLAINED_ERRORS + 1))
|
|
;;
|
|
*) REDEPLOY_UNEXPLAINED_ERRORS=$((REDEPLOY_UNEXPLAINED_ERRORS + 1)) ;;
|
|
esac
|
|
;;
|
|
*'AMQP connection recovered; cleared held replies for fresh redelivery'*)
|
|
if [ "$pending_inbox" -gt 0 ]; then
|
|
pending_inbox=$((pending_inbox - 1))
|
|
REDEPLOY_RECOVERED_AMQP_ERRORS=$((REDEPLOY_RECOVERED_AMQP_ERRORS + 1))
|
|
fi
|
|
;;
|
|
*'AMQP lead mailbox connection recovered; cleared held messages for fresh redelivery'*)
|
|
if [ "$pending_lead_mailbox" -gt 0 ]; then
|
|
pending_lead_mailbox=$((pending_lead_mailbox - 1))
|
|
REDEPLOY_RECOVERED_AMQP_ERRORS=$((REDEPLOY_RECOVERED_AMQP_ERRORS + 1))
|
|
fi
|
|
;;
|
|
esac
|
|
done < "$log_file"
|
|
|
|
REDEPLOY_UNEXPLAINED_ERRORS=$((REDEPLOY_UNEXPLAINED_ERRORS + pending_inbox + pending_lead_mailbox))
|
|
}
|
|
|
|
# fleetd #512 part 2 — the negative check. #493's failure (an uncaught exception in a shutdown
|
|
# thread) never passes through the logger: the JVM's default uncaught-exception handler prints
|
|
# straight to stderr, so the line never carries a level, so classify_amqp_connection_errors's
|
|
# ERROR/SEVERE token scan is structurally blind to it — measured on two hosts, including one where
|
|
# even a syslog PRIORITY filter is blind to it too (fd 1 and fd 2 collapse to one socket there, so
|
|
# every uncaught-exception line lands at priority 6/info). The fix is to grep the shape instead of
|
|
# the level: `Exception in thread` at the start of a line (the handler's own banner) or
|
|
# `NoClassDefFoundError` anywhere in it (the one real instance seen so far, but not the only shape
|
|
# this could take). Kept as its own function, never folded into classify_amqp_connection_errors —
|
|
# this is not an AMQP concern, and the two must stay independently readable and independently
|
|
# testable.
|
|
#
|
|
# Sets REDEPLOY_UNCAUGHT_EXCEPTION_COUNT (lines matched) and REDEPLOY_UNCAUGHT_EXCEPTION_SAMPLE
|
|
# (the first matching line, "" if none) so a caller can report both a count and a concrete quote
|
|
# without re-reading the file. Pure: reads $1, sets globals, no side effects.
|
|
scan_uncaught_exceptions() {
|
|
local log_file="$1" line
|
|
REDEPLOY_UNCAUGHT_EXCEPTION_COUNT=0
|
|
REDEPLOY_UNCAUGHT_EXCEPTION_SAMPLE=""
|
|
while IFS= read -r line || [ -n "$line" ]; do
|
|
case "$line" in
|
|
'Exception in thread'*|*NoClassDefFoundError*)
|
|
REDEPLOY_UNCAUGHT_EXCEPTION_COUNT=$((REDEPLOY_UNCAUGHT_EXCEPTION_COUNT + 1))
|
|
[ -n "$REDEPLOY_UNCAUGHT_EXCEPTION_SAMPLE" ] || REDEPLOY_UNCAUGHT_EXCEPTION_SAMPLE="$line"
|
|
;;
|
|
esac
|
|
done < "$log_file"
|
|
}
|
|
|
|
# fleetd #512 part 2 — the positive check. fleetd #522 added a `log.info` at the very end of
|
|
# SessionManager.drainAll's normal path (never in a `finally` — see the ticket discussion for why
|
|
# that distinction matters): "drain complete: released=N abandoned=M (still BUSY at the shutdown
|
|
# deadline)", printed once, on every successful drain, including the all-zero case. A drain that
|
|
# dies partway through never reaches that statement, so the line's ABSENCE is a real signal — unlike
|
|
# the ERROR-count check above, this one does not depend on the failure happening to throw.
|
|
#
|
|
# Sets REDEPLOY_DRAIN_COMPLETE_LINE to the matching line (last one, though drainAll runs at most
|
|
# once per shutdown so there should never be more than one) or "" if absent. Pure, same shape as
|
|
# scan_uncaught_exceptions above.
|
|
find_drain_complete_line() {
|
|
local log_file="$1"
|
|
REDEPLOY_DRAIN_COMPLETE_LINE="$(grep -F 'drain complete: released=' "$log_file" | tail -1 || true)"
|
|
}
|
|
|
|
# fleetd #512 part 2 — THE TRAP, and the reason this is one function instead of two independent
|
|
# checks the caller ORs together. Absence of the drain-complete line has TWO causes that need
|
|
# OPPOSITE handling, and a naive "line absent -> the drain died" reading collapses them exactly the
|
|
# way this whole ticket exists to stop: the line is emitted by the daemon being STOPPED, which is
|
|
# running the OLD jar. Until a redeploy has landed fleetd #522 once, every previous daemon predates
|
|
# the line and cannot emit it no matter how cleanly it drained — so on the very first redeploy after
|
|
# #522 merged, "absent" means "too old to know how", not "died". Only once BOTH signals — this
|
|
# line's absence AND scan_uncaught_exceptions' result — have been read together can the three real
|
|
# outcomes be told apart:
|
|
#
|
|
# complete -> the line is present: the drain finished. Name the counts it reported.
|
|
# died -> the line is absent AND an uncaught-exception shape was found: the drain died. Name
|
|
# what was found.
|
|
# unknown -> the line is absent AND no exception shape either: cannot tell. Say so, and say why
|
|
# (predates the line, or failed without throwing) — never worded as a pass or a
|
|
# failure, and never reassuring: "ok, no ERROR lines" one level up is the exact mistake
|
|
# this ticket exists to fix, and this outcome must not reproduce it.
|
|
#
|
|
# A fourth case, n/a, covers a cold start or a "loaded but wasn't running" restart: no previous
|
|
# daemon was actually stopped THIS run, so there is no shutdown window in $log_file to have an
|
|
# opinion about at all — scanning it anyway would read the NEW daemon's own startup lines and could
|
|
# misreport "cannot tell" on every clean cold start. had_previous_daemon carries that fact in from
|
|
# the caller (it already knows $OLD_PID) rather than this function re-deriving it from log content.
|
|
#
|
|
# Same shape as swap_if_built/refuse_drain_gate (fleetd #521/#528): the decision (which of the four
|
|
# outcomes applies) and the action (which ok/warn line to print, and setting REDEPLOY_DRAIN_STATE
|
|
# for the "result" section below to consult) live together in ONE function that the main flow calls
|
|
# unconditionally — there is no guard left in the main flow to remove, invert, or bypass
|
|
# independently of this function. Never calls die(): #512's own decision is to warn loudly and let
|
|
# the redeploy stand, because by the time this is detectable the new daemon is already up and
|
|
# healthy and failing here would give the operator nothing to do differently.
|
|
report_shutdown_drain() {
|
|
local log_file="$1" had_previous_daemon="$2"
|
|
REDEPLOY_UNCAUGHT_EXCEPTION_COUNT=0
|
|
REDEPLOY_UNCAUGHT_EXCEPTION_SAMPLE=""
|
|
REDEPLOY_DRAIN_COMPLETE_LINE=""
|
|
|
|
if [ "$had_previous_daemon" != 1 ]; then
|
|
REDEPLOY_DRAIN_STATE="n/a"
|
|
ok "no previous daemon was running before this restart — nothing to check for a died shutdown drain"
|
|
return 0
|
|
fi
|
|
|
|
find_drain_complete_line "$log_file"
|
|
scan_uncaught_exceptions "$log_file"
|
|
|
|
if [ -n "$REDEPLOY_DRAIN_COMPLETE_LINE" ]; then
|
|
REDEPLOY_DRAIN_STATE="complete"
|
|
ok "previous daemon's shutdown drain finished: $REDEPLOY_DRAIN_COMPLETE_LINE"
|
|
elif [ "$REDEPLOY_UNCAUGHT_EXCEPTION_COUNT" -gt 0 ]; then
|
|
REDEPLOY_DRAIN_STATE="died"
|
|
warn "previous daemon's shutdown drain DIED — no drain-complete line, and an uncaught exception"
|
|
warn "was found in its shutdown window ($REDEPLOY_UNCAUGHT_EXCEPTION_COUNT line(s)):"
|
|
warn " $REDEPLOY_UNCAUGHT_EXCEPTION_SAMPLE"
|
|
warn "Some sessions from the PREVIOUS daemon may not have been released."
|
|
else
|
|
REDEPLOY_DRAIN_STATE="unknown"
|
|
warn "cannot tell whether the previous daemon's shutdown drain finished — no drain-complete line"
|
|
warn "and no uncaught-exception shape either. This is NOT a pass and NOT a failure: it means"
|
|
warn "either that daemon predates fleetd #522's drain-complete log line, or its drain failed"
|
|
warn "without throwing (hung, or returned early)."
|
|
fi
|
|
}
|
|
|
|
# fleetd #517: extracted so the suite can call this decision directly, the same way #510 extracted
|
|
# wait_for_daemon_exit so its ordering became checkable. Before this, the only test of the drain-gate
|
|
# abort message was a grep of this script's own source for the wording — so mutating the `if` below
|
|
# to `if false` (making the branch unreachable) left every test green, because the wording was still
|
|
# sitting in the file. Pure: only decides which message applies and prints it, no side effects, so a
|
|
# test can call it directly with an in-memory staged path instead of driving the real drain-gate flow
|
|
# (which needs a live $OLD_PID and an interactive prompt neither test can supply).
|
|
#
|
|
# The four cases:
|
|
# build ran, staged jar present -> names the staged jar and how to finish or discard it
|
|
# build ran, staged jar absent -> "nothing changed" (nothing was staged this run either)
|
|
# --no-build, staged jar present -> ALSO "nothing changed", deliberately: --no-build itself builds
|
|
# and stages nothing (see require_no_build_jar above), so a staged jar found here is a leftover
|
|
# from an earlier, unrelated run. THIS run truly changed nothing, and the next DO_BUILD=1 run
|
|
# wipes that leftover before it builds (`rm -f "$JAR_STAGED"` in the build section above) — so
|
|
# there is nothing here for the operator to lose track of.
|
|
# --no-build, staged jar absent -> "nothing changed"
|
|
drain_gate_refusal() {
|
|
local do_build="$1" staged_path="$2"
|
|
if [ "$do_build" = 1 ] && [ -f "$staged_path" ]; then
|
|
printf 'aborted — the running daemon was NOT touched, but the freshly built jar is sitting at
|
|
%s, not yet swapped into %s. Rerun WITHOUT --no-build to finish the restart —
|
|
the freshly built jar is no longer at the live path that --no-build requires — or
|
|
remove %s by hand if you want to discard this build.' "$staged_path" "$JAR" "$staged_path"
|
|
else
|
|
printf 'aborted — nothing changed'
|
|
fi
|
|
}
|
|
|
|
# fleetd #528 — drain_gate_refusal above is well tested (four cases, all direct), but nothing made
|
|
# the MAIN FLOW's abort actually consult it. Before this, the main flow read
|
|
# `die "$(drain_gate_refusal "$DO_BUILD" "$JAR_STAGED")"` directly, and mutating that one line to a
|
|
# flat `die "aborted — nothing changed"` left the whole suite at exit 0 with zero FAIL lines and
|
|
# byte-identical output to a clean run — every one of drain_gate_refusal's own tests still passed,
|
|
# because they call the predicate directly and never touch this call site. That silently reinstated
|
|
# the exact defect #517 was filed to fix. Same shape as #521/#526's should_swap/swap_if_built: a
|
|
# predicate alone is not enough, because a test proving the predicate is right cannot also prove the
|
|
# main flow consults it. So the decision (drain_gate_refusal) and the action (die) now live together
|
|
# in ONE function, and the main flow calls it unconditionally instead of building the die() call
|
|
# itself — there is no guard left in the main flow to remove, invert, or bypass independently of this
|
|
# function. drain_gate_refusal stays separate and separately tested because the message-selection
|
|
# logic is worth naming and testing on its own; refuse_drain_gate is the only thing that ever dies.
|
|
#
|
|
# What the behavioural tests above still cannot pin on their own: deleting the call to this function
|
|
# from the main flow altogether — they call refuse_drain_gate directly, never through the main flow,
|
|
# because sourcing stops before the main flow ever runs (see the SOURCED guard below). That gap is
|
|
# closed the same way swap_if_built's is: test_refuse_drain_gate_call_site_present greps this script
|
|
# for the real invocation, the same shape test_swap_ordered_after_wait_and_before_start already uses
|
|
# for the swap call. Deliberately NOT written out here as a literal quoted string, so this comment
|
|
# itself can never become a second match for that test's needle.
|
|
refuse_drain_gate() {
|
|
local do_build="$1" staged_path="$2"
|
|
die "$(drain_gate_refusal "$do_build" "$staged_path")"
|
|
}
|
|
|
|
# CB-600: sourceable for testing. When this file is SOURCED (not executed) it stops here — nothing
|
|
# below runs — so a test harness can `source` it to call check_log_path_matches_plist (or the
|
|
# other pure helpers above) against a throwaway plist fixture without ever reaching the mutating
|
|
# flow (build/stop/start) or touching the real daemon or launchd. On a normal `./redeploy-fleetd.sh`
|
|
# invocation `(return 0 2>/dev/null)` fails (return is illegal at top level of an executed script),
|
|
# so this whole block is a no-op and every line below still runs exactly as before.
|
|
if (return 0 2>/dev/null); then
|
|
return 0
|
|
fi
|
|
|
|
# ---------------------------------------------------------------- report state
|
|
|
|
say "current state"
|
|
OLD_PID="$(running_pid)"
|
|
if [ -n "$OLD_PID" ]; then
|
|
ok "daemon running, pid $OLD_PID"
|
|
else
|
|
warn "no daemon running — this will be a cold start"
|
|
fi
|
|
ok "jar on disk: $(jar_id) ($([ -f "$JAR" ] && date -r "$JAR" '+%Y-%m-%d %H:%M:%S' || echo 'none'))"
|
|
ok "HEAD: $(git -C "$REPO" log --oneline -1)"
|
|
|
|
# CB-594 / fleetd #492: supervision state. Installed and loaded are different facts — a
|
|
# copied-but-never-loaded plist (or an unloaded systemd unit) supervises nothing, and a loaded
|
|
# label/unit with no file backing it is still what its supervisor will act on.
|
|
if launchd_installed; then
|
|
ok "launchd agent installed: $LAUNCHD_PLIST"
|
|
else
|
|
warn "launchd agent NOT installed."
|
|
fi
|
|
if systemd_installed; then
|
|
ok "systemd --user unit installed: $SYSTEMD_UNIT"
|
|
else
|
|
warn "systemd --user unit NOT installed ($SYSTEMD_UNIT)."
|
|
fi
|
|
|
|
# fleetd #492: decide which of the two (if either) actually supervises this daemon, and refuse
|
|
# outright — before touching anything — if that cannot be told apart (see require_drivable_
|
|
# supervisor above). --check reaches this same line, so a host with an undrivable supervisor is
|
|
# reported as a failure even in --check, without ever reaching the build/stop/start steps.
|
|
# fleetd #492 follow-up: detect_supervisor runs as $(...), so only the printed line survives —
|
|
# unpack kind and detail from it HERE, in this shell, before calling anything downstream. See the
|
|
# constraints comment above detect_supervisor for why this cannot be done any other way.
|
|
SUPERVISOR_RAW="$(detect_supervisor)"
|
|
SUPERVISOR_KIND="${SUPERVISOR_RAW%%"$SUPERVISOR_DETAIL_SEP"*}"
|
|
SUPERVISOR_UNCLEAR_DETAIL="${SUPERVISOR_RAW#*"$SUPERVISOR_DETAIL_SEP"}"
|
|
require_drivable_supervisor "$SUPERVISOR_KIND"
|
|
ok "supervisor detected: $SUPERVISOR_KIND"
|
|
SUPERVISED=0
|
|
case "$SUPERVISOR_KIND" in
|
|
launchd)
|
|
SUPERVISED=1
|
|
ok "launchd agent loaded ($LAUNCHD_LABEL) — launchd supervises this daemon"
|
|
# CB-600: fail loudly here, before ANY other check runs, if this script and the loaded plist
|
|
# would read different log files — every check after this point is worthless otherwise.
|
|
check_log_path_matches_plist "$OUT" "$LAUNCHD_PLIST"
|
|
;;
|
|
systemd)
|
|
SUPERVISED=1
|
|
ok "systemd --user unit active ($SYSTEMD_UNIT) — systemd supervises this daemon"
|
|
;;
|
|
none)
|
|
warn "no supervisor loaded — this script is the only thing that will restart the daemon."
|
|
;;
|
|
*)
|
|
# fleetd #492 follow-up: this block only DISPLAYS state, it changes nothing yet — so a value
|
|
# it doesn't recognise gets reported, not an abort that goes silent on exactly the state most
|
|
# worth seeing. (Unreachable today: require_drivable_supervisor above already died on
|
|
# "ambiguous"/"unclear" before this case runs. Guards the value nobody has invented yet.)
|
|
warn "unrecognised supervisor kind: '$SUPERVISOR_KIND' — detect_supervisor returned a value this block does not know; continuing to report the rest of the state."
|
|
;;
|
|
esac
|
|
|
|
# The trap with no log line. Checked in a LOGIN shell, because that is how the daemon is started
|
|
# below. Never prints the value — only whether it resolved.
|
|
if zsh -lc '[ -n "${WORKER_GITEA_TOKEN:-}" ]' 2>/dev/null; then
|
|
ok "WORKER_GITEA_TOKEN resolves in a login shell"
|
|
else
|
|
warn "WORKER_GITEA_TOKEN is EMPTY in a login shell."
|
|
warn "The daemon will start fine and workers will silently fail to open PRs."
|
|
warn "Fix \${SHARED_ENV}/tools/secrets.sh before relying on worker checkpoints."
|
|
fi
|
|
|
|
# Same trap, second variable (CB-591). A profile's `tokenEnv:` is resolved from the DAEMON's own
|
|
# process environment by HerdrPeerLauncher.resolveEnv, so a token added to secrets.sh after the
|
|
# daemon started is simply absent. The launcher then injects an empty token and llm.ltms.dev answers
|
|
# 401 — long after the restart, and with nothing tying the two together.
|
|
if zsh -lc '[ -n "${AI_GATEWAY_TOKEN:-}" ]' 2>/dev/null; then
|
|
ok "AI_GATEWAY_TOKEN resolves in a login shell"
|
|
else
|
|
warn "AI_GATEWAY_TOKEN is EMPTY in a login shell."
|
|
warn "Any profile whose tokenEnv is AI_GATEWAY_TOKEN will get an empty token and 401 at the gateway."
|
|
warn "This only matters once a profile points at llm.ltms.dev — harmless before that."
|
|
fi
|
|
|
|
# Third variable, same trap (CB-635). broker.uriEnv names the env var holding the AMQP URI, so the
|
|
# password stays out of fleetd.yaml — but that moves the failure into the environment. If the
|
|
# variable is empty the daemon still starts: since #152 it warns and falls back to the in-memory
|
|
# reply inbox, so nothing crashes and replies simply stop surviving a restart. Only this check says
|
|
# so before the fact. Read the name out of the config so a renamed key cannot make the check lie.
|
|
BROKER_URI_ENV=$(sed -n 's/^[[:space:]]*uriEnv:[[:space:]]*\([A-Za-z_][A-Za-z0-9_]*\).*/\1/p' "$MODULE/fleetd.yaml" | head -1)
|
|
if [ -z "$BROKER_URI_ENV" ]; then
|
|
ok "no broker.uriEnv configured — reply inbox is in-memory by design"
|
|
elif zsh -lc "[ -n \"\${$BROKER_URI_ENV:-}\" ]" 2>/dev/null; then
|
|
ok "$BROKER_URI_ENV (broker.uriEnv) resolves in a login shell"
|
|
else
|
|
warn "$BROKER_URI_ENV (broker.uriEnv) is EMPTY in a login shell."
|
|
warn "The daemon will start and fall back to the IN-MEMORY reply inbox."
|
|
warn "Replies stop surviving a restart — a held report is lost, not delayed."
|
|
fi
|
|
|
|
if [ "$CHECK_ONLY" = 1 ]; then
|
|
say "--check: nothing changed"
|
|
exit 0
|
|
fi
|
|
|
|
# ---------------------------------------------------------------------- build
|
|
# Deliberately before the stop: a failed build must never leave the fleet down.
|
|
|
|
if [ "$DO_BUILD" = 1 ]; then
|
|
say "build"
|
|
# fleetd #493: wipe a leftover staged jar from a previous failed/interrupted run BEFORE doing
|
|
# anything else, so that run's leftovers can never be mistaken for this run's output.
|
|
rm -f "$JAR_STAGED"
|
|
BUILD_LOG="$(mktemp -t fleetd-build.XXXXXX)"
|
|
echo " log: $BUILD_LOG"
|
|
if ! mvn -f "$MODULE/pom.xml" clean install > "$BUILD_LOG" 2>&1; then
|
|
grep -E 'ERROR|BUILD FAILURE|Tests run:.*Failures: [1-9]|Tests run:.*Errors: [1-9]' "$BUILD_LOG" \
|
|
| head -20 || true
|
|
die "build failed — the running daemon was NOT touched. Full log: $BUILD_LOG"
|
|
fi
|
|
grep -E '^\[INFO\] Tests run:.*Failures' "$BUILD_LOG" | tail -1 | sed 's/^\[INFO\] / /' || true
|
|
ok "BUILD SUCCESS"
|
|
# fleetd #493: move the freshly built jar off the live path immediately — the running (OLD)
|
|
# daemon, if any, is still up at this point (build always runs before stop). From here until the
|
|
# swap step below (after the OLD daemon is confirmed gone), $JAR_STAGED is the only artefact this
|
|
# script treats as "the new jar" — $JAR itself is not touched again until the swap.
|
|
stage_built_jar
|
|
ok "jar now: $(jar_id "$JAR_STAGED")"
|
|
else
|
|
say "build skipped (--no-build)"
|
|
# fleetd #493: --no-build never builds or stages anything — it restarts whatever jar is already
|
|
# sitting at the live path from an earlier successful run. Same check, same message as before.
|
|
require_no_build_jar
|
|
fi
|
|
|
|
# ----------------------------------------------------------------- drain gate
|
|
|
|
if [ -n "$OLD_PID" ] && [ "$ASSUME_YES" = 0 ]; then
|
|
say "drain check"
|
|
echo " A restart drops every in-flight ticket and rendezvous. A member's report"
|
|
echo " is NOT recoverable once its ticket is gone."
|
|
echo
|
|
echo " Confirm with fleet_list that no members are live, and fleet_poll anything"
|
|
echo " you still want, BEFORE continuing."
|
|
echo
|
|
read -r -p " Fleet drained? type yes to restart: " reply
|
|
if [ "$reply" != "yes" ]; then
|
|
# fleetd #493 / #517 / #528: "nothing changed" would be a lie once a build has run and staged a
|
|
# jar — see drain_gate_refusal above for the full decision and why each of its four cases reads
|
|
# the way it does. refuse_drain_gate composes that message AND calls die itself, so this guard
|
|
# has nothing left of its own to get wrong beyond whether it calls refuse_drain_gate at all.
|
|
refuse_drain_gate "$DO_BUILD" "$JAR_STAGED"
|
|
fi
|
|
fi
|
|
|
|
# ------------------------------------------------------------------ stop
|
|
#
|
|
# CB-594 / fleetd #492: when SUPERVISED, the supervisor owns the stop — never a raw `kill` here. A
|
|
# bare SIGTERM makes this JVM exit 143 even with its shutdown hook running to completion (verified
|
|
# separately: a throwaway Java process with an equivalent shutdown hook, sent SIGTERM from a login
|
|
# shell that could `wait` on it directly, reported exit code 143 every time — never 0). launchd's
|
|
# KeepAlive.SuccessfulExit=false and systemd's Restart=on-failure both treat any nonzero exit as a
|
|
# crash and restart the OLD jar, which would race this script's own restart of the NEW one.
|
|
# `launchctl unload` avoids that race by deregistering the job first, so no KeepAlive is left
|
|
# armed when the process actually stops. `systemctl --user stop` needs no such dance: unlike
|
|
# KeepAlive, systemd's Restart= does not fire on a deliberate stop, only on an unexpected exit of
|
|
# an active unit.
|
|
|
|
if [ -n "$OLD_PID" ]; then
|
|
say "stop"
|
|
RESTART_MARK="$(wc -l < "$OUT" 2>/dev/null || echo 0)" # verify a FRESH line appears later
|
|
case "$SUPERVISOR_KIND" in
|
|
launchd)
|
|
echo " supervision is ON (launchd): using 'launchctl unload' (not kill) so launchd's own"
|
|
echo " KeepAlive cannot restart the OLD jar out from under this script — see the CB-594"
|
|
echo " comment above."
|
|
launchctl unload -w "$LAUNCHD_PLIST" \
|
|
|| die "launchctl unload failed — the daemon may still be under supervision; investigate before retrying"
|
|
;;
|
|
systemd)
|
|
echo " supervision is ON (systemd --user): using 'systemctl --user stop' (not kill) so"
|
|
echo " systemd's own Restart=on-failure cannot restart the OLD jar out from under this"
|
|
echo " script — see the fleetd #492 comment above."
|
|
systemctl --user stop "$SYSTEMD_UNIT" \
|
|
|| die "'systemctl --user stop $SYSTEMD_UNIT' failed — the daemon may still be under supervision; investigate before retrying"
|
|
;;
|
|
none)
|
|
kill "$OLD_PID"
|
|
;;
|
|
*)
|
|
# fleetd #492 follow-up: this block ACTS (stops the daemon one specific way per kind) — an
|
|
# unrecognised value must never fall through to a default action, silently picking the wrong
|
|
# one (or none at all) while reporting success. (Unreachable today: require_drivable_
|
|
# supervisor already died before this runs. Guards the value nobody has invented yet.)
|
|
die "detect_supervisor returned an unrecognized value '$SUPERVISOR_KIND' at the stop step —
|
|
refusing to guess how to stop a daemon under an unknown supervisor. The daemon was NOT
|
|
stopped." ;;
|
|
esac
|
|
if ! wait_for_daemon_exit "$STOP_WAIT"; then
|
|
die "pid $OLD_PID still alive after ${STOP_WAIT}s. Not escalating to kill -9 automatically:
|
|
the shutdown hook releases sessions and worktrees in order, and killing it hard can
|
|
leave worktrees and panes behind. Investigate, then kill -9 by hand if you accept that."
|
|
fi
|
|
ok "pid $OLD_PID exited"
|
|
elif [ "$SUPERVISOR_KIND" = "launchd" ]; then
|
|
# Loaded but not currently running (e.g. throttled after a crash loop). Unload it anyway so the
|
|
# start step below does a clean load, never a load stacked on an already-loaded label.
|
|
# fleetd #504: unload_launchd_if_loaded (above) tolerates a genuine already-unloaded answer but
|
|
# dies on a real `launchctl` failure — never a bare `|| true` that would print `ok` either way.
|
|
say "stop"
|
|
RESTART_MARK="$(wc -l < "$OUT" 2>/dev/null || echo 0)"
|
|
unload_launchd_if_loaded
|
|
ok "launchd agent unloaded (was already not running)"
|
|
elif [ "$SUPERVISOR_KIND" = "systemd" ]; then
|
|
# Same case for systemd: the unit is known/active-capable but not currently running. `stop` on an
|
|
# already-stopped unit is a harmless no-op — kept for symmetry with the launchd branch above.
|
|
# fleetd #504: stop_systemd_if_loaded (above) tolerates that genuine no-op but dies on a real
|
|
# `systemctl` failure — never a bare `|| true` that would print `ok` either way.
|
|
say "stop"
|
|
RESTART_MARK="$(wc -l < "$OUT" 2>/dev/null || echo 0)"
|
|
stop_systemd_if_loaded
|
|
ok "systemd --user unit stopped (was already not running)"
|
|
else
|
|
RESTART_MARK="$(wc -l < "$OUT" 2>/dev/null || echo 0)"
|
|
fi
|
|
|
|
# ------------------------------------------------------------------ swap
|
|
#
|
|
# fleetd #493: every branch above has now either confirmed the OLD daemon actually exited
|
|
# (wait_for_daemon_exit, above) or established there was never one running to begin with. Only
|
|
# NOW is it safe to put the freshly built jar at the path the NEXT `java -jar` (direct, or via
|
|
# launchd/systemd's ExecStart) will read from — this mv is the one and only write to $JAR anywhere
|
|
# in this script's mutating flow. If it fails, do not start: die() below exits before "start" runs.
|
|
swap_if_built "$DO_BUILD"
|
|
|
|
# ------------------------------------------------------------------ start
|
|
# Unsupervised: login shell (zsh -l) is what puts the secrets on the daemon's environment, and cwd
|
|
# must be fleetd/ because the daemon resolves fleetd.yaml, logs/ and target/ relative to it.
|
|
# Supervised (launchd): launchd does both — deploy/dev.ltms.fleetd.plist points ProgramArguments at
|
|
# scripts/fleetd-launchd-wrapper.sh (CB-594), which is what execs the login shell in launchd's
|
|
# place, and WorkingDirectory in the plist already pins fleetd/.
|
|
# Supervised (systemd --user): the unit does both too — measured on the second host, ExecStart is
|
|
# `/bin/zsh -lc "exec java -jar target/fleetd.jar fleetd.yaml"` (a login shell, same reason as
|
|
# above) and WorkingDirectory is already pinned to fleetd/.
|
|
|
|
say "start"
|
|
case "$SUPERVISOR_KIND" in
|
|
launchd)
|
|
echo " supervision is ON (launchd): using 'launchctl load' so launchd starts and keeps"
|
|
echo " supervising this process, instead of a manual nohup that launchd would know nothing"
|
|
echo " about."
|
|
# CB-600: 'launchctl unload -w' above already persisted Disabled=true for this label. A load -w
|
|
# that succeeds clears it; a load -w that FAILS leaves the agent both stopped and disabled — worse
|
|
# than before this script ran, because a later reboot or login will not bring it back either. One
|
|
# retry covers a transient race (e.g. launchd not yet fully done deregistering); if it still fails,
|
|
# die with the exact recovery command rather than a bare "failed".
|
|
if ! launchctl load -w "$LAUNCHD_PLIST" 2>/dev/null; then
|
|
warn "launchctl load failed on the first attempt — retrying once after a short pause"
|
|
sleep 2
|
|
launchctl load -w "$LAUNCHD_PLIST" || die "launchctl load failed twice.
|
|
The agent is now STOPPED and DISABLED — it will NOT come back on its own, not even after a
|
|
reboot or login, because 'launchctl unload -w' above persisted Disabled=true and load -w
|
|
never got the chance to clear it. Recover with:
|
|
launchctl load -w \"$LAUNCHD_PLIST\"
|
|
If that still fails, check 'launchctl list $LAUNCHD_LABEL', validate the plist with
|
|
'plutil -lint \"$LAUNCHD_PLIST\"', and check $OUT before assuming a retry will succeed."
|
|
fi
|
|
;;
|
|
systemd)
|
|
echo " supervision is ON (systemd --user): using 'systemctl --user start' so systemd starts"
|
|
echo " and keeps supervising this process, instead of a manual nohup it would know nothing"
|
|
echo " about."
|
|
systemctl --user start "$SYSTEMD_UNIT" || die "'systemctl --user start $SYSTEMD_UNIT' failed.
|
|
Check 'systemctl --user status $SYSTEMD_UNIT' and $OUT before assuming a retry will succeed."
|
|
;;
|
|
none)
|
|
# Absolute jar path so `ps` names which checkout is running.
|
|
( cd "$MODULE" && zsh -lc "nohup java -jar '$JAR' >> fleetd.out 2>&1 &" )
|
|
;;
|
|
*)
|
|
# fleetd #492 follow-up: this block ACTS (starts the daemon one specific way per kind) — an
|
|
# unrecognised value must never fall through to a default action, silently picking the wrong
|
|
# one (or none at all) while reporting success. (Unreachable today: require_drivable_
|
|
# supervisor already died before this runs. Guards the value nobody has invented yet.)
|
|
die "detect_supervisor returned an unrecognized value '$SUPERVISOR_KIND' at the start step —
|
|
refusing to guess how to start a daemon under an unknown supervisor. The daemon was NOT
|
|
started." ;;
|
|
esac
|
|
|
|
for _ in $(seq 10); do
|
|
NEW_PID="$(running_pid)"
|
|
[ -n "$NEW_PID" ] && break
|
|
sleep 1
|
|
done
|
|
[ -n "${NEW_PID:-}" ] || die "no process appeared. Last lines of $OUT:
|
|
$(tail -20 "$OUT" 2>/dev/null)"
|
|
[ "$NEW_PID" != "${OLD_PID:-}" ] || die "pid unchanged ($NEW_PID) — the old daemon never died"
|
|
ok "started, pid $NEW_PID"
|
|
|
|
# ------------------------------------------------------------------ verify
|
|
|
|
say "verify"
|
|
|
|
HEALTH_BODY=""
|
|
for _ in $(seq "$HEALTH_WAIT"); do
|
|
if HEALTH_BODY="$(curl -fsS --max-time 2 "$HEALTH" 2>/dev/null)"; then break; fi
|
|
HEALTH_BODY=""
|
|
sleep 1
|
|
done
|
|
|
|
if [ -z "$HEALTH_BODY" ]; then
|
|
# 503 still means the daemon is up — it means herdr is unreachable. Say which.
|
|
CODE="$(curl -s -o /dev/null -w '%{http_code}' --max-time 2 "$HEALTH" 2>/dev/null || echo 000)"
|
|
if [ "$CODE" = "503" ]; then
|
|
warn "/healthz answers 503 degraded — the daemon is up but herdr is unreachable."
|
|
warn "Spawns will fail. Check herdr before delegating anything."
|
|
curl -s --max-time 2 "$HEALTH" 2>/dev/null | head -3 || true
|
|
else
|
|
die "/healthz never answered within ${HEALTH_WAIT}s (last code: $CODE). Last lines of $OUT:
|
|
$(tail -30 "$OUT" 2>/dev/null)"
|
|
fi
|
|
else
|
|
ok "/healthz 200 — $HEALTH_BODY"
|
|
warn "healthz green only proves herdr ANSWERS. If its protocol number changed, spawns can still"
|
|
warn "fail — prove a real spawn before trusting the fleet."
|
|
fi
|
|
|
|
# A fresh listening line, strictly after the restart mark. An old daemon that never died would
|
|
# otherwise let an old line pass for a new one.
|
|
if tail -n "+$((RESTART_MARK + 1))" "$OUT" 2>/dev/null | grep -q 'fleetd listening'; then
|
|
ok "$(tail -n "+$((RESTART_MARK + 1))" "$OUT" | grep 'fleetd listening' | tail -1)"
|
|
else
|
|
warn "no fresh 'fleetd listening' line after the restart — check $OUT yourself"
|
|
fi
|
|
|
|
# Config keys the daemon accepted or deferred at boot. This is usually WHY you restarted.
|
|
say "config at boot"
|
|
tail -n "+$((RESTART_MARK + 1))" "$OUT" 2>/dev/null \
|
|
| grep -iE 'deferred|classification:|fleet health:|coverage' | tail -8 | sed 's/^/ /' \
|
|
|| echo " (nothing reported)"
|
|
|
|
# Errors since the restart, anchored to the marker so old noise cannot leak in. Keep the fresh
|
|
# region in a file because the classifier must preserve the order of errors and recoveries.
|
|
FRESH_LOG="$(mktemp -t fleetd-fresh-log.XXXXXX)"
|
|
trap 'rm -f "$FRESH_LOG"' EXIT
|
|
tail -n "+$((RESTART_MARK + 1))" "$OUT" > "$FRESH_LOG" 2>/dev/null || true
|
|
classify_amqp_connection_errors "$FRESH_LOG"
|
|
|
|
# fleetd #512 part 2: the previous daemon's shutdown drain, checked in the same fresh-log region —
|
|
# see report_shutdown_drain above for the full decision (four outcomes, one of them a deliberate
|
|
# "cannot tell"). HAD_OLD_PID crosses in whether a previous daemon was actually stopped this run;
|
|
# see the function's own comment for why that matters.
|
|
say "previous daemon's shutdown drain"
|
|
HAD_OLD_PID=0; [ -n "$OLD_PID" ] && HAD_OLD_PID=1
|
|
report_shutdown_drain "$FRESH_LOG" "$HAD_OLD_PID"
|
|
|
|
# fleetd #492: checked here, after healthz and the fresh-log check have both had time to run, so a
|
|
# supervisor that revives the OLD jar a few seconds late is caught too. Every check above (healthz
|
|
# 200, jar id, the fresh 'listening' line) is satisfied by EITHER daemon if two are alive — this is
|
|
# the only one that can tell.
|
|
assert_single_daemon "$(running_pid)"
|
|
|
|
say "result"
|
|
ok "pid $NEW_PID, jar $(jar_id)"
|
|
if [ "$REDEPLOY_ERROR_COUNT" -eq 0 ]; then
|
|
# fleetd #512 item 4: this line must not print when the shutdown-drain check above found the
|
|
# previous daemon's drain died, or could not tell — either would make "no ERROR lines" read as a
|
|
# clean bill of health it is not (an uncaught exception never carries an ERROR token to begin
|
|
# with, so this count alone cannot see that failure). "complete" and "n/a" are the only two
|
|
# outcomes report_shutdown_drain sets that mean nothing is wrong there.
|
|
if [ "$REDEPLOY_DRAIN_STATE" = "complete" ] || [ "$REDEPLOY_DRAIN_STATE" = "n/a" ]; then
|
|
ok "no ERROR lines since restart"
|
|
fi
|
|
elif [ "$REDEPLOY_UNEXPLAINED_ERRORS" -eq 0 ]; then
|
|
ok "$REDEPLOY_RECOVERED_AMQP_ERRORS AMQP connection reset ERROR lines recovered since restart"
|
|
else
|
|
warn "$REDEPLOY_ERROR_COUNT ERROR lines since restart:"
|
|
grep -E ' (ERROR|SEVERE) ' "$FRESH_LOG" | tail -5 | sed 's/^/ /'
|
|
fi
|
|
echo
|
|
echo " Next: call fleet_whoami and confirm it still answers 'primary'. A lead whose tab label"
|
|
echo " no longer matches fleet.leaders.*.tab is demoted to worker and refuses orchestration."
|
|
echo
|