CB-594: make supervision and a working fleet possible at the same time
Adds scripts/bridged-launchd-wrapper.sh so the launchd-run daemon still gets WORKER_GITEA_TOKEN/AI_GATEWAY_TOKEN by execing through a login shell (launchd never sources secrets.sh itself). bridged now logs at startup which required token env vars (derived from each profile's tokenEnv/gitTokenEnv, not a hand-written list) resolved or are MISSING, by name only. Fills in the real paths in deploy/dev.ltms.bridged.plist for this host and points it at the wrapper. scripts/redeploy-bridged.sh now detects a loaded launchd agent and uses launchctl unload/load instead of a raw kill+nohup, because a bare SIGTERM exits this JVM at 143 (measured) which KeepAlive.SuccessfulExit=false reads as a crash and would race the script's own restart; --check reports installed/loaded state and stays read-only.
This commit is contained in:
Executable
+32
@@ -0,0 +1,32 @@
|
||||
#!/usr/bin/env bash
|
||||
#
|
||||
# CB-594 — the only reason this file exists: launchd does not run a login shell.
|
||||
#
|
||||
# WORKER_GITEA_TOKEN and AI_GATEWAY_TOKEN live in ${SHARED_ENV}/tools/secrets.sh, sourced only by a
|
||||
# LOGIN shell (.zprofile/.zshrc etc). launchd execs a job's ProgramArguments directly — no shell, no
|
||||
# profile, nothing sourced (the plist's own PATH comment documents the same gap one variable over).
|
||||
# A daemon started that way boots fine and looks healthy; the failure is invisible until a worker
|
||||
# tries to open a PR (WORKER_GITEA_TOKEN empty) or a gateway profile gets a 401 (AI_GATEWAY_TOKEN
|
||||
# empty) — hours later, with nothing tying the two together (CB-591, CLAUDE.md "Redeploying the
|
||||
# daemon"). Bridged now also logs which required secret names resolved at startup (see
|
||||
# Bridged.reportRequiredSecrets), but that log line can only tell the truth if the tokens had a
|
||||
# chance to be sourced in the first place — which is this script's entire job.
|
||||
#
|
||||
# So: launchd execs THIS script instead of java directly. This script execs a login shell
|
||||
# ('zsh -l'), which sources secrets.sh, and that shell execs the real command in its place — one
|
||||
# process throughout (exec, not a subshell fork), so launchd's PID tracking, KeepAlive, and
|
||||
# StandardOut/ErrorPath all still see the one process they expect.
|
||||
#
|
||||
# The plist passes the full command as THIS script's own arguments, e.g.:
|
||||
# ProgramArguments = [ .../bridged-launchd-wrapper.sh, /path/to/java, -jar, /path/to/bridged.jar,
|
||||
# bridged.yaml ]
|
||||
# so the wrapper stays generic and the actual command lives in exactly one place (the plist), not
|
||||
# duplicated here.
|
||||
set -euo pipefail
|
||||
|
||||
if [ "$#" -eq 0 ]; then
|
||||
echo "bridged-launchd-wrapper.sh: no command given — check the plist's ProgramArguments" >&2
|
||||
exit 2
|
||||
fi
|
||||
|
||||
exec /bin/zsh -lc 'exec "$@"' -- "$@"
|
||||
@@ -5,13 +5,14 @@
|
||||
# A merge is not a deployment: the running daemon holds the jar it was started with, so code merged
|
||||
# to main does nothing until this runs. See CLAUDE.md -> "Redeploying the daemon".
|
||||
#
|
||||
# This script exists to turn five remembered traps into one auditable command:
|
||||
# This script exists to turn six remembered traps into one auditable command:
|
||||
#
|
||||
# 1. A piped `mvn` hides BUILD FAILURE behind a zero exit, so the build here is never piped.
|
||||
# 2. The daemon must start from a LOGIN shell, or the tokens it hands to members are empty:
|
||||
# WORKER_GITEA_TOKEN (workers cannot open a PR) and AI_GATEWAY_TOKEN (401 at llm.ltms.dev).
|
||||
# Both are read from the DAEMON's own environment at spawn time, so a value added to
|
||||
# secrets.sh after startup is absent. Nothing logs this, so the script checks and says so.
|
||||
# secrets.sh after startup is absent. Nothing logs this here, so the script checks and says
|
||||
# so — and since CB-594, bridged's own startup log says so too, by env var name.
|
||||
# 3. An old daemon that never actually died looks identical from the outside, so the script waits
|
||||
# for the process to exit and for the port to free before it starts a new one.
|
||||
# 4. "It started" is not "it works": the script polls /healthz until it answers, and reports the
|
||||
@@ -19,6 +20,13 @@
|
||||
# mismatch.
|
||||
# 5. Restarting under live members drops their tickets, so the script refuses unless you confirm
|
||||
# the fleet is drained.
|
||||
# 6. CB-594 — the launchd agent (deploy/dev.ltms.bridged.plist), if installed and loaded, is a
|
||||
# SECOND supervisor: its KeepAlive.SuccessfulExit=false restarts the daemon on any nonzero
|
||||
# exit, and a bare SIGTERM makes this JVM exit 143 even with its shutdown hook running to
|
||||
# completion (measured — see the CB-594 report). A plain `kill` here would race launchd's own
|
||||
# restart of the OLD jar. So this script detects whether the agent is loaded and, only then,
|
||||
# swaps `kill` + manual `nohup` for `launchctl unload`/`load` — the one supervisor in control
|
||||
# at any moment is whichever one you asked to act, never both.
|
||||
#
|
||||
# Usage:
|
||||
# scripts/redeploy-bridged.sh # build, confirm, restart, verify
|
||||
@@ -43,13 +51,17 @@ HEALTH='http://127.0.0.1:8765/healthz'
|
||||
STOP_WAIT=30 # seconds to wait for a clean exit before reporting failure
|
||||
HEALTH_WAIT=60 # seconds to wait for /healthz to answer after start
|
||||
|
||||
# CB-594: the launchd agent this script must not fight with (see trap 6 above).
|
||||
LAUNCHD_LABEL='dev.ltms.bridged'
|
||||
LAUNCHD_PLIST="$HOME/Library/LaunchAgents/$LAUNCHD_LABEL.plist"
|
||||
|
||||
DO_BUILD=1; ASSUME_YES=0; CHECK_ONLY=0
|
||||
for arg in "$@"; do
|
||||
case "$arg" in
|
||||
--yes|-y) ASSUME_YES=1 ;;
|
||||
--no-build) DO_BUILD=0 ;;
|
||||
--check) CHECK_ONLY=1 ;;
|
||||
-h|--help) sed -n '3,30p' "${BASH_SOURCE[0]}"; exit 0 ;;
|
||||
-h|--help) sed -n '3,37p' "${BASH_SOURCE[0]}"; exit 0 ;;
|
||||
*) echo "unknown option: $arg (try --help)" >&2; exit 2 ;;
|
||||
esac
|
||||
done
|
||||
@@ -61,6 +73,11 @@ die() { printf '\n FAIL %s\n\n' "$*" >&2; exit 1; }
|
||||
|
||||
jar_id() { [ -f "$JAR" ] && shasum -a 256 "$JAR" | cut -c1-12 || echo "absent"; }
|
||||
running_pid() { pgrep -f "$PATTERN" || true; }
|
||||
# `launchctl list <label>` exits 0 iff the label is loaded (registered with launchd) — true whether
|
||||
# or not it is currently running, which is exactly "supervision is active" for our purposes. Read-
|
||||
# only: neither helper below changes anything, so both are also safe under --check.
|
||||
launchd_installed() { [ -f "$LAUNCHD_PLIST" ]; }
|
||||
launchd_loaded() { launchctl list "$LAUNCHD_LABEL" >/dev/null 2>&1; }
|
||||
|
||||
# ---------------------------------------------------------------- report state
|
||||
|
||||
@@ -74,6 +91,22 @@ fi
|
||||
ok "jar on disk: $(jar_id) ($([ -f "$JAR" ] && date -r "$JAR" '+%Y-%m-%d %H:%M:%S' || echo 'none'))"
|
||||
ok "HEAD: $(git -C "$REPO" log --oneline -1)"
|
||||
|
||||
# CB-594: supervision state. Installed and loaded are different facts — a copied-but-never-loaded
|
||||
# plist supervises nothing, and a loaded label with no file backing it (rare, but possible after an
|
||||
# edited/moved plist) is still what launchd will act on.
|
||||
if launchd_installed; then
|
||||
ok "launchd agent installed: $LAUNCHD_PLIST"
|
||||
else
|
||||
warn "launchd agent NOT installed (no supervision — a crash will not restart the daemon)."
|
||||
fi
|
||||
SUPERVISED=0
|
||||
if launchd_loaded; then
|
||||
SUPERVISED=1
|
||||
ok "launchd agent loaded ($LAUNCHD_LABEL) — launchd supervises this daemon"
|
||||
else
|
||||
warn "launchd agent not loaded — this script is the only thing that will restart the daemon."
|
||||
fi
|
||||
|
||||
# The trap with no log line. Checked in a LOGIN shell, because that is how the daemon is started
|
||||
# below. Never prints the value — only whether it resolved.
|
||||
if zsh -lc '[ -n "${WORKER_GITEA_TOKEN:-}" ]' 2>/dev/null; then
|
||||
@@ -137,11 +170,26 @@ if [ -n "$OLD_PID" ] && [ "$ASSUME_YES" = 0 ]; then
|
||||
fi
|
||||
|
||||
# ------------------------------------------------------------------ stop
|
||||
#
|
||||
# CB-594: when SUPERVISED, launchd owns the stop — never a raw `kill` here. A bare SIGTERM makes
|
||||
# this JVM exit 143 even with its shutdown hook running to completion (verified separately: a
|
||||
# throwaway Java process with an equivalent shutdown hook, sent SIGTERM from a login shell that
|
||||
# could `wait` on it directly, reported exit code 143 every time — never 0). launchd's
|
||||
# KeepAlive.SuccessfulExit=false treats any nonzero exit as a crash and restarts the OLD jar,
|
||||
# which would race this script's own restart of the NEW one. `launchctl unload` avoids that race
|
||||
# by deregistering the job first, so no KeepAlive is left armed when the process actually stops.
|
||||
|
||||
if [ -n "$OLD_PID" ]; then
|
||||
say "stop"
|
||||
RESTART_MARK="$(wc -l < "$OUT" 2>/dev/null || echo 0)" # verify a FRESH line appears later
|
||||
kill "$OLD_PID"
|
||||
if [ "$SUPERVISED" = 1 ]; then
|
||||
echo " supervision is ON: using 'launchctl unload' (not kill) so launchd's own KeepAlive"
|
||||
echo " cannot restart the OLD jar out from under this script — see the CB-594 comment above."
|
||||
launchctl unload -w "$LAUNCHD_PLIST" \
|
||||
|| die "launchctl unload failed — the daemon may still be under supervision; investigate before retrying"
|
||||
else
|
||||
kill "$OLD_PID"
|
||||
fi
|
||||
for _ in $(seq "$STOP_WAIT"); do
|
||||
[ -z "$(running_pid)" ] && break
|
||||
sleep 1
|
||||
@@ -152,18 +200,33 @@ if [ -n "$OLD_PID" ]; then
|
||||
leave worktrees and panes behind. Investigate, then kill -9 by hand if you accept that."
|
||||
fi
|
||||
ok "pid $OLD_PID exited"
|
||||
elif [ "$SUPERVISED" = 1 ]; then
|
||||
# Loaded but not currently running (e.g. throttled after a crash loop). Unload it anyway so the
|
||||
# start step below does a clean load, never a load stacked on an already-loaded label.
|
||||
say "stop"
|
||||
RESTART_MARK="$(wc -l < "$OUT" 2>/dev/null || echo 0)"
|
||||
launchctl unload -w "$LAUNCHD_PLIST" 2>/dev/null || true
|
||||
ok "launchd agent unloaded (was already not running)"
|
||||
else
|
||||
RESTART_MARK="$(wc -l < "$OUT" 2>/dev/null || echo 0)"
|
||||
fi
|
||||
|
||||
# ------------------------------------------------------------------ start
|
||||
# Login shell (zsh -l) is what puts the secrets on the daemon's environment. cwd must be bridged/
|
||||
# because the daemon resolves bridged.yaml, logs/ and target/ relative to it.
|
||||
# Unsupervised: login shell (zsh -l) is what puts the secrets on the daemon's environment, and cwd
|
||||
# must be bridged/ because the daemon resolves bridged.yaml, logs/ and target/ relative to it.
|
||||
# Supervised: launchd does both — deploy/dev.ltms.bridged.plist points ProgramArguments at
|
||||
# scripts/bridged-launchd-wrapper.sh (CB-594), which is what execs the login shell in launchd's
|
||||
# place, and WorkingDirectory in the plist already pins bridged/.
|
||||
|
||||
say "start"
|
||||
# Absolute jar path so `ps` names which checkout is running; cwd still bridged/ because the daemon
|
||||
# resolves bridged.yaml, logs/ and target/ relative to it.
|
||||
( cd "$BRIDGED" && zsh -lc "nohup java -jar '$JAR' >> bridged.out 2>&1 &" )
|
||||
if [ "$SUPERVISED" = 1 ]; then
|
||||
echo " supervision is ON: using 'launchctl load' so launchd starts and keeps supervising this"
|
||||
echo " process, instead of a manual nohup that launchd would know nothing about."
|
||||
launchctl load -w "$LAUNCHD_PLIST" || die "launchctl load failed"
|
||||
else
|
||||
# Absolute jar path so `ps` names which checkout is running.
|
||||
( cd "$BRIDGED" && zsh -lc "nohup java -jar '$JAR' >> bridged.out 2>&1 &" )
|
||||
fi
|
||||
|
||||
for _ in $(seq 10); do
|
||||
NEW_PID="$(running_pid)"
|
||||
|
||||
Reference in New Issue
Block a user