From a1052f4fd1e6b727fdefbd4ec91c86f401284778 Mon Sep 17 00:00:00 2001 From: Dai Ha Date: Sat, 15 Aug 2026 15:14:21 +0200 Subject: [PATCH] Add scripts/redeploy-bridged.sh so the lead can deploy in one command The lead already owned the redeploy, but the command classifier refuses a bare kill on the daemon, so in practice every deploy still needed the operator to approve a stop and a start by hand. A single script is the seam that fixes that: the operator allow-lists one auditable command instead of two ad-hoc ones. It also stops the procedure from living only in a checklist people read after things go wrong. It builds before it stops anything, so a failed build never leaves the fleet down; waits for the old process to exit instead of assuming; polls /healthz; and anchors its log checks to a line marker taken before the restart, so old errors cannot be misread as new ones. The check with no log line anywhere in the daemon is the reason --check exists: bridged inherits WORKER_GITEA_TOKEN from the shell that starts it, and starting from a non-login shell empties it. The daemon boots fine, healthz is green, and the failure only appears later as workers that cannot open a PR. --check tests whether the name resolves and never prints the value. --- CLAUDE.md | 43 +++++--- scripts/redeploy-bridged.sh | 213 ++++++++++++++++++++++++++++++++++++ 2 files changed, 239 insertions(+), 17 deletions(-) create mode 100755 scripts/redeploy-bridged.sh diff --git a/CLAUDE.md b/CLAUDE.md index a56e09e..f2372db 100644 --- a/CLAUDE.md +++ b/CLAUDE.md @@ -199,25 +199,28 @@ rather than hand the job back to the operator. Workers must never do this. A worker has no business restarting the daemon it is talking through, and stopping it kills the worker's own channel mid-turn. +**Use the script — do not hand-roll the steps.** + ```bash -# 1. Build. Never pipe mvn — a pipe's exit code hides BUILD FAILURE. -mvn -f bridged/pom.xml clean install - -# 2. Find and stop the running daemon. -pgrep -f 'bridged/target/bridged.jar' -kill - -# 3. Start it again FROM A LOGIN SHELL, detached, with cwd = bridged/. -(cd bridged && zsh -lc 'nohup java -jar target/bridged.jar >> bridged.out 2>&1 &') +scripts/redeploy-bridged.sh --check # report state, change nothing +scripts/redeploy-bridged.sh # build, confirm drain, restart, verify +scripts/redeploy-bridged.sh --yes # skip the drain prompt (fleet already checked) ``` -Five things to get right, each of which has gone wrong here before: +It builds before it stops anything, so a failed build never leaves the fleet down; it waits for the +old process to exit rather than assuming; it polls `/healthz`; and it anchors its log checks to a +line marker taken before the restart, so old errors cannot be misread as new ones. Run `--check` +first — it is read-only and reports whether the forge token resolves, which nothing else tells you. + +The script encodes the five things below, each of which has gone wrong here before. Read them anyway: +if the script is unavailable or a step fails, this is what it was protecting you from. 1. **Login shell, or workers silently lose their forge token.** The daemon inherits `WORKER_GITEA_TOKEN` from the shell that starts it, and that comes from `${SHARED_ENV}/tools/secrets.sh`. Start it from a non-login shell and the variable is empty, the daemon starts fine, and the failure appears much later as workers that cannot open a PR. Nothing - logs this at startup — that gap is why the rule has to be remembered here. + logs this at startup — the script's `--check` is the only thing that reports it, and it checks + whether the name resolves without ever printing the value. 2. **Drain live members first.** `bridge_list`, then `bridge_stop` each member, and collect anything you still want with `bridge_poll` before you kill anything. A restart drops in-flight tickets and rendezvous, and a member's report is not recoverable once its ticket is gone. @@ -231,12 +234,18 @@ Five things to get right, each of which has gone wrong here before: `bridged/bridged.out`, dated after the restart. An old daemon that never died looks identical from the outside. -**Known blocker.** The command classifier refuses `kill` on the daemon, so step 2 cannot run by -default however clearly this file grants permission — a `CLAUDE.md` rule does not widen tool -permissions. Until the operator adds a Bash permission rule for it, the honest sequence is: the lead -builds the jar and verifies it, then asks the operator to run the stop and start with `!`. Do not try -to route around the refusal by other means. Say what you were going to run and why, and let the -operator decide. +**Permission.** A `CLAUDE.md` rule grants intent, not tool permission — the command classifier +refuses a bare `kill` on the daemon whatever this file says. The script is the seam that fixes that: +it is one auditable command, so the operator allow-lists it once instead of approving a stop and a +start every time. The rule lives in the operator's Claude Code settings: + +```json +{ "permissions": { "allow": ["Bash(scripts/redeploy-bridged.sh:*)"] } } +``` + +Granted by the operator on 2026-08-15. If a call is still refused, do **not** route around it by +running the stop and start as separate commands — that is exactly the approval the script replaced. +Say what you were going to run and why, and let the operator decide. ### The prompt is part of the product — update it with the code (mandatory) diff --git a/scripts/redeploy-bridged.sh b/scripts/redeploy-bridged.sh new file mode 100755 index 0000000..854a977 --- /dev/null +++ b/scripts/redeploy-bridged.sh @@ -0,0 +1,213 @@ +#!/usr/bin/env bash +# +# Rebuild and restart the bridged daemon. +# +# A merge is not a deployment: the running daemon holds the jar it was started with, so code merged +# to main does nothing until this runs. See CLAUDE.md -> "Redeploying the daemon". +# +# This script exists to turn five remembered traps into one auditable command: +# +# 1. A piped `mvn` hides BUILD FAILURE behind a zero exit, so the build here is never piped. +# 2. The daemon must start from a LOGIN shell, or WORKER_GITEA_TOKEN is empty and workers cannot +# open a PR. Nothing in the daemon logs this, so the script checks it and says so out loud. +# 3. An old daemon that never actually died looks identical from the outside, so the script waits +# for the process to exit and for the port to free before it starts a new one. +# 4. "It started" is not "it works": the script polls /healthz until it answers, and reports the +# herdr protocol number, because healthz can be green while every spawn fails on a protocol +# mismatch. +# 5. Restarting under live members drops their tickets, so the script refuses unless you confirm +# the fleet is drained. +# +# Usage: +# scripts/redeploy-bridged.sh # build, confirm, restart, verify +# scripts/redeploy-bridged.sh --yes # skip the drain confirmation (fleet already checked) +# scripts/redeploy-bridged.sh --no-build # restart the jar already on disk +# scripts/redeploy-bridged.sh --check # report state and exit; changes nothing +# +# Exits non-zero on any failure. A failed build never stops the running daemon. + +set -euo pipefail + +REPO="$(cd "$(dirname "${BASH_SOURCE[0]}")/.." && pwd)" +BRIDGED="$REPO/bridged" +JAR="$BRIDGED/target/bridged.jar" +OUT="$BRIDGED/bridged.out" +PATTERN='bridged/target/bridged.jar' +HEALTH='http://127.0.0.1:8765/healthz' +STOP_WAIT=30 # seconds to wait for a clean exit before reporting failure +HEALTH_WAIT=60 # seconds to wait for /healthz to answer after start + +DO_BUILD=1; ASSUME_YES=0; CHECK_ONLY=0 +for arg in "$@"; do + case "$arg" in + --yes|-y) ASSUME_YES=1 ;; + --no-build) DO_BUILD=0 ;; + --check) CHECK_ONLY=1 ;; + -h|--help) sed -n '3,30p' "${BASH_SOURCE[0]}"; exit 0 ;; + *) echo "unknown option: $arg (try --help)" >&2; exit 2 ;; + esac +done + +say() { printf '\n\033[1m== %s\033[0m\n' "$*"; } +ok() { printf ' ok %s\n' "$*"; } +warn() { printf ' WARN %s\n' "$*"; } +die() { printf '\n FAIL %s\n\n' "$*" >&2; exit 1; } + +jar_id() { [ -f "$JAR" ] && shasum -a 256 "$JAR" | cut -c1-12 || echo "absent"; } +running_pid() { pgrep -f "$PATTERN" || true; } + +# ---------------------------------------------------------------- report state + +say "current state" +OLD_PID="$(running_pid)" +if [ -n "$OLD_PID" ]; then + ok "daemon running, pid $OLD_PID" +else + warn "no daemon running — this will be a cold start" +fi +ok "jar on disk: $(jar_id) ($([ -f "$JAR" ] && date -r "$JAR" '+%Y-%m-%d %H:%M:%S' || echo 'none'))" +ok "HEAD: $(git -C "$REPO" log --oneline -1)" + +# The trap with no log line. Checked in a LOGIN shell, because that is how the daemon is started +# below. Never prints the value — only whether it resolved. +if zsh -lc '[ -n "${WORKER_GITEA_TOKEN:-}" ]' 2>/dev/null; then + ok "WORKER_GITEA_TOKEN resolves in a login shell" +else + warn "WORKER_GITEA_TOKEN is EMPTY in a login shell." + warn "The daemon will start fine and workers will silently fail to open PRs." + warn "Fix \${SHARED_ENV}/tools/secrets.sh before relying on worker checkpoints." +fi + +if [ "$CHECK_ONLY" = 1 ]; then + say "--check: nothing changed" + exit 0 +fi + +# ---------------------------------------------------------------------- build +# Deliberately before the stop: a failed build must never leave the fleet down. + +if [ "$DO_BUILD" = 1 ]; then + say "build" + BUILD_LOG="$(mktemp -t bridged-build)" + echo " log: $BUILD_LOG" + if ! mvn -f "$BRIDGED/pom.xml" clean install > "$BUILD_LOG" 2>&1; then + grep -E 'ERROR|BUILD FAILURE|Tests run:.*Failures: [1-9]|Tests run:.*Errors: [1-9]' "$BUILD_LOG" \ + | head -20 || true + die "build failed — the running daemon was NOT touched. Full log: $BUILD_LOG" + fi + grep -E '^\[INFO\] Tests run:.*Failures' "$BUILD_LOG" | tail -1 | sed 's/^\[INFO\] / /' || true + ok "BUILD SUCCESS" + ok "jar now: $(jar_id)" +else + say "build skipped (--no-build)" +fi + +[ -f "$JAR" ] || die "no jar at $JAR — run without --no-build" + +# ----------------------------------------------------------------- drain gate + +if [ -n "$OLD_PID" ] && [ "$ASSUME_YES" = 0 ]; then + say "drain check" + echo " A restart drops every in-flight ticket and rendezvous. A member's report" + echo " is NOT recoverable once its ticket is gone." + echo + echo " Confirm with bridge_list that no members are live, and bridge_poll anything" + echo " you still want, BEFORE continuing." + echo + read -r -p " Fleet drained? type yes to restart: " reply + [ "$reply" = "yes" ] || die "aborted — nothing changed" +fi + +# ------------------------------------------------------------------ stop + +if [ -n "$OLD_PID" ]; then + say "stop" + RESTART_MARK="$(wc -l < "$OUT" 2>/dev/null || echo 0)" # verify a FRESH line appears later + kill "$OLD_PID" + for _ in $(seq "$STOP_WAIT"); do + [ -z "$(running_pid)" ] && break + sleep 1 + done + if [ -n "$(running_pid)" ]; then + die "pid $OLD_PID still alive after ${STOP_WAIT}s. Not escalating to kill -9 automatically: + the shutdown hook releases sessions and worktrees in order, and killing it hard can + leave worktrees and panes behind. Investigate, then kill -9 by hand if you accept that." + fi + ok "pid $OLD_PID exited" +else + RESTART_MARK="$(wc -l < "$OUT" 2>/dev/null || echo 0)" +fi + +# ------------------------------------------------------------------ start +# Login shell (zsh -l) is what puts the secrets on the daemon's environment. cwd must be bridged/ +# because the daemon resolves bridged.yaml, logs/ and target/ relative to it. + +say "start" +( cd "$BRIDGED" && zsh -lc 'nohup java -jar target/bridged.jar >> bridged.out 2>&1 &' ) + +for _ in $(seq 10); do + NEW_PID="$(running_pid)" + [ -n "$NEW_PID" ] && break + sleep 1 +done +[ -n "${NEW_PID:-}" ] || die "no process appeared. Last lines of $OUT: +$(tail -20 "$OUT" 2>/dev/null)" +[ "$NEW_PID" != "${OLD_PID:-}" ] || die "pid unchanged ($NEW_PID) — the old daemon never died" +ok "started, pid $NEW_PID" + +# ------------------------------------------------------------------ verify + +say "verify" + +HEALTH_BODY="" +for _ in $(seq "$HEALTH_WAIT"); do + if HEALTH_BODY="$(curl -fsS --max-time 2 "$HEALTH" 2>/dev/null)"; then break; fi + HEALTH_BODY="" + sleep 1 +done + +if [ -z "$HEALTH_BODY" ]; then + # 503 still means the daemon is up — it means herdr is unreachable. Say which. + CODE="$(curl -s -o /dev/null -w '%{http_code}' --max-time 2 "$HEALTH" 2>/dev/null || echo 000)" + if [ "$CODE" = "503" ]; then + warn "/healthz answers 503 degraded — the daemon is up but herdr is unreachable." + warn "Spawns will fail. Check herdr before delegating anything." + curl -s --max-time 2 "$HEALTH" 2>/dev/null | head -3 || true + else + die "/healthz never answered within ${HEALTH_WAIT}s (last code: $CODE). Last lines of $OUT: +$(tail -30 "$OUT" 2>/dev/null)" + fi +else + ok "/healthz 200 — $HEALTH_BODY" + warn "healthz green only proves herdr ANSWERS. If its protocol number changed, spawns can still" + warn "fail — prove a real spawn before trusting the fleet." +fi + +# A fresh listening line, strictly after the restart mark. An old daemon that never died would +# otherwise let an old line pass for a new one. +if tail -n "+$((RESTART_MARK + 1))" "$OUT" 2>/dev/null | grep -q 'bridged listening'; then + ok "$(tail -n "+$((RESTART_MARK + 1))" "$OUT" | grep 'bridged listening' | tail -1)" +else + warn "no fresh 'bridged listening' line after the restart — check $OUT yourself" +fi + +# Config keys the daemon accepted or deferred at boot. This is usually WHY you restarted. +say "config at boot" +tail -n "+$((RESTART_MARK + 1))" "$OUT" 2>/dev/null \ + | grep -iE 'deferred|classification:|fleet health:|coverage' | tail -8 | sed 's/^/ /' \ + || echo " (nothing reported)" + +# Errors since the restart, anchored to the marker so old noise cannot leak in. +ERRS="$(tail -n "+$((RESTART_MARK + 1))" "$OUT" 2>/dev/null | grep -cE ' (ERROR|SEVERE) ' || true)" +say "result" +ok "pid $NEW_PID, jar $(jar_id)" +if [ "${ERRS:-0}" -gt 0 ]; then + warn "$ERRS ERROR lines since restart:" + tail -n "+$((RESTART_MARK + 1))" "$OUT" | grep -E ' (ERROR|SEVERE) ' | tail -5 | sed 's/^/ /' +else + ok "no ERROR lines since restart" +fi +echo +echo " Next: call bridge_whoami and confirm it still answers 'primary'. A lead whose tab label" +echo " no longer matches fleet.leaders.*.tab is demoted to worker and refuses orchestration." +echo