Add scripts/redeploy-bridged.sh so the lead can deploy in one command
CI / build (push) Successful in 57s
CI / contract (push) Successful in 1m23s

The lead already owned the redeploy, but the command classifier refuses a bare
kill on the daemon, so in practice every deploy still needed the operator to
approve a stop and a start by hand. A single script is the seam that fixes
that: the operator allow-lists one auditable command instead of two ad-hoc
ones.

It also stops the procedure from living only in a checklist people read after
things go wrong. It builds before it stops anything, so a failed build never
leaves the fleet down; waits for the old process to exit instead of assuming;
polls /healthz; and anchors its log checks to a line marker taken before the
restart, so old errors cannot be misread as new ones.

The check with no log line anywhere in the daemon is the reason --check exists:
bridged inherits WORKER_GITEA_TOKEN from the shell that starts it, and starting
from a non-login shell empties it. The daemon boots fine, healthz is green, and
the failure only appears later as workers that cannot open a PR. --check tests
whether the name resolves and never prints the value.
This commit is contained in:
Dai Ha
2026-08-15 15:14:21 +02:00
parent 8c9904a7c4
commit a1052f4fd1
2 changed files with 239 additions and 17 deletions
+26 -17
View File
@@ -199,25 +199,28 @@ rather than hand the job back to the operator.
Workers must never do this. A worker has no business restarting the daemon it is talking through,
and stopping it kills the worker's own channel mid-turn.
**Use the script — do not hand-roll the steps.**
```bash
# 1. Build. Never pipe mvn — a pipe's exit code hides BUILD FAILURE.
mvn -f bridged/pom.xml clean install
# 2. Find and stop the running daemon.
pgrep -f 'bridged/target/bridged.jar'
kill <pid>
# 3. Start it again FROM A LOGIN SHELL, detached, with cwd = bridged/.
(cd bridged && zsh -lc 'nohup java -jar target/bridged.jar >> bridged.out 2>&1 &')
scripts/redeploy-bridged.sh --check # report state, change nothing
scripts/redeploy-bridged.sh # build, confirm drain, restart, verify
scripts/redeploy-bridged.sh --yes # skip the drain prompt (fleet already checked)
```
Five things to get right, each of which has gone wrong here before:
It builds before it stops anything, so a failed build never leaves the fleet down; it waits for the
old process to exit rather than assuming; it polls `/healthz`; and it anchors its log checks to a
line marker taken before the restart, so old errors cannot be misread as new ones. Run `--check`
first — it is read-only and reports whether the forge token resolves, which nothing else tells you.
The script encodes the five things below, each of which has gone wrong here before. Read them anyway:
if the script is unavailable or a step fails, this is what it was protecting you from.
1. **Login shell, or workers silently lose their forge token.** The daemon inherits
`WORKER_GITEA_TOKEN` from the shell that starts it, and that comes from
`${SHARED_ENV}/tools/secrets.sh`. Start it from a non-login shell and the variable is empty, the
daemon starts fine, and the failure appears much later as workers that cannot open a PR. Nothing
logs this at startup — that gap is why the rule has to be remembered here.
logs this at startup — the script's `--check` is the only thing that reports it, and it checks
whether the name resolves without ever printing the value.
2. **Drain live members first.** `bridge_list`, then `bridge_stop` each member, and collect anything
you still want with `bridge_poll` before you kill anything. A restart drops in-flight tickets and
rendezvous, and a member's report is not recoverable once its ticket is gone.
@@ -231,12 +234,18 @@ Five things to get right, each of which has gone wrong here before:
`bridged/bridged.out`, dated after the restart. An old daemon that never died looks identical from
the outside.
**Known blocker.** The command classifier refuses `kill` on the daemon, so step 2 cannot run by
default however clearly this file grants permission — a `CLAUDE.md` rule does not widen tool
permissions. Until the operator adds a Bash permission rule for it, the honest sequence is: the lead
builds the jar and verifies it, then asks the operator to run the stop and start with `!`. Do not try
to route around the refusal by other means. Say what you were going to run and why, and let the
operator decide.
**Permission.** A `CLAUDE.md` rule grants intent, not tool permission — the command classifier
refuses a bare `kill` on the daemon whatever this file says. The script is the seam that fixes that:
it is one auditable command, so the operator allow-lists it once instead of approving a stop and a
start every time. The rule lives in the operator's Claude Code settings:
```json
{ "permissions": { "allow": ["Bash(scripts/redeploy-bridged.sh:*)"] } }
```
Granted by the operator on 2026-08-15. If a call is still refused, do **not** route around it by
running the stop and start as separate commands — that is exactly the approval the script replaced.
Say what you were going to run and why, and let the operator decide.
### The prompt is part of the product — update it with the code (mandatory)
+213
View File
@@ -0,0 +1,213 @@
#!/usr/bin/env bash
#
# Rebuild and restart the bridged daemon.
#
# A merge is not a deployment: the running daemon holds the jar it was started with, so code merged
# to main does nothing until this runs. See CLAUDE.md -> "Redeploying the daemon".
#
# This script exists to turn five remembered traps into one auditable command:
#
# 1. A piped `mvn` hides BUILD FAILURE behind a zero exit, so the build here is never piped.
# 2. The daemon must start from a LOGIN shell, or WORKER_GITEA_TOKEN is empty and workers cannot
# open a PR. Nothing in the daemon logs this, so the script checks it and says so out loud.
# 3. An old daemon that never actually died looks identical from the outside, so the script waits
# for the process to exit and for the port to free before it starts a new one.
# 4. "It started" is not "it works": the script polls /healthz until it answers, and reports the
# herdr protocol number, because healthz can be green while every spawn fails on a protocol
# mismatch.
# 5. Restarting under live members drops their tickets, so the script refuses unless you confirm
# the fleet is drained.
#
# Usage:
# scripts/redeploy-bridged.sh # build, confirm, restart, verify
# scripts/redeploy-bridged.sh --yes # skip the drain confirmation (fleet already checked)
# scripts/redeploy-bridged.sh --no-build # restart the jar already on disk
# scripts/redeploy-bridged.sh --check # report state and exit; changes nothing
#
# Exits non-zero on any failure. A failed build never stops the running daemon.
set -euo pipefail
REPO="$(cd "$(dirname "${BASH_SOURCE[0]}")/.." && pwd)"
BRIDGED="$REPO/bridged"
JAR="$BRIDGED/target/bridged.jar"
OUT="$BRIDGED/bridged.out"
PATTERN='bridged/target/bridged.jar'
HEALTH='http://127.0.0.1:8765/healthz'
STOP_WAIT=30 # seconds to wait for a clean exit before reporting failure
HEALTH_WAIT=60 # seconds to wait for /healthz to answer after start
DO_BUILD=1; ASSUME_YES=0; CHECK_ONLY=0
for arg in "$@"; do
case "$arg" in
--yes|-y) ASSUME_YES=1 ;;
--no-build) DO_BUILD=0 ;;
--check) CHECK_ONLY=1 ;;
-h|--help) sed -n '3,30p' "${BASH_SOURCE[0]}"; exit 0 ;;
*) echo "unknown option: $arg (try --help)" >&2; exit 2 ;;
esac
done
say() { printf '\n\033[1m== %s\033[0m\n' "$*"; }
ok() { printf ' ok %s\n' "$*"; }
warn() { printf ' WARN %s\n' "$*"; }
die() { printf '\n FAIL %s\n\n' "$*" >&2; exit 1; }
jar_id() { [ -f "$JAR" ] && shasum -a 256 "$JAR" | cut -c1-12 || echo "absent"; }
running_pid() { pgrep -f "$PATTERN" || true; }
# ---------------------------------------------------------------- report state
say "current state"
OLD_PID="$(running_pid)"
if [ -n "$OLD_PID" ]; then
ok "daemon running, pid $OLD_PID"
else
warn "no daemon running — this will be a cold start"
fi
ok "jar on disk: $(jar_id) ($([ -f "$JAR" ] && date -r "$JAR" '+%Y-%m-%d %H:%M:%S' || echo 'none'))"
ok "HEAD: $(git -C "$REPO" log --oneline -1)"
# The trap with no log line. Checked in a LOGIN shell, because that is how the daemon is started
# below. Never prints the value — only whether it resolved.
if zsh -lc '[ -n "${WORKER_GITEA_TOKEN:-}" ]' 2>/dev/null; then
ok "WORKER_GITEA_TOKEN resolves in a login shell"
else
warn "WORKER_GITEA_TOKEN is EMPTY in a login shell."
warn "The daemon will start fine and workers will silently fail to open PRs."
warn "Fix \${SHARED_ENV}/tools/secrets.sh before relying on worker checkpoints."
fi
if [ "$CHECK_ONLY" = 1 ]; then
say "--check: nothing changed"
exit 0
fi
# ---------------------------------------------------------------------- build
# Deliberately before the stop: a failed build must never leave the fleet down.
if [ "$DO_BUILD" = 1 ]; then
say "build"
BUILD_LOG="$(mktemp -t bridged-build)"
echo " log: $BUILD_LOG"
if ! mvn -f "$BRIDGED/pom.xml" clean install > "$BUILD_LOG" 2>&1; then
grep -E 'ERROR|BUILD FAILURE|Tests run:.*Failures: [1-9]|Tests run:.*Errors: [1-9]' "$BUILD_LOG" \
| head -20 || true
die "build failed — the running daemon was NOT touched. Full log: $BUILD_LOG"
fi
grep -E '^\[INFO\] Tests run:.*Failures' "$BUILD_LOG" | tail -1 | sed 's/^\[INFO\] / /' || true
ok "BUILD SUCCESS"
ok "jar now: $(jar_id)"
else
say "build skipped (--no-build)"
fi
[ -f "$JAR" ] || die "no jar at $JAR — run without --no-build"
# ----------------------------------------------------------------- drain gate
if [ -n "$OLD_PID" ] && [ "$ASSUME_YES" = 0 ]; then
say "drain check"
echo " A restart drops every in-flight ticket and rendezvous. A member's report"
echo " is NOT recoverable once its ticket is gone."
echo
echo " Confirm with bridge_list that no members are live, and bridge_poll anything"
echo " you still want, BEFORE continuing."
echo
read -r -p " Fleet drained? type yes to restart: " reply
[ "$reply" = "yes" ] || die "aborted — nothing changed"
fi
# ------------------------------------------------------------------ stop
if [ -n "$OLD_PID" ]; then
say "stop"
RESTART_MARK="$(wc -l < "$OUT" 2>/dev/null || echo 0)" # verify a FRESH line appears later
kill "$OLD_PID"
for _ in $(seq "$STOP_WAIT"); do
[ -z "$(running_pid)" ] && break
sleep 1
done
if [ -n "$(running_pid)" ]; then
die "pid $OLD_PID still alive after ${STOP_WAIT}s. Not escalating to kill -9 automatically:
the shutdown hook releases sessions and worktrees in order, and killing it hard can
leave worktrees and panes behind. Investigate, then kill -9 by hand if you accept that."
fi
ok "pid $OLD_PID exited"
else
RESTART_MARK="$(wc -l < "$OUT" 2>/dev/null || echo 0)"
fi
# ------------------------------------------------------------------ start
# Login shell (zsh -l) is what puts the secrets on the daemon's environment. cwd must be bridged/
# because the daemon resolves bridged.yaml, logs/ and target/ relative to it.
say "start"
( cd "$BRIDGED" && zsh -lc 'nohup java -jar target/bridged.jar >> bridged.out 2>&1 &' )
for _ in $(seq 10); do
NEW_PID="$(running_pid)"
[ -n "$NEW_PID" ] && break
sleep 1
done
[ -n "${NEW_PID:-}" ] || die "no process appeared. Last lines of $OUT:
$(tail -20 "$OUT" 2>/dev/null)"
[ "$NEW_PID" != "${OLD_PID:-}" ] || die "pid unchanged ($NEW_PID) — the old daemon never died"
ok "started, pid $NEW_PID"
# ------------------------------------------------------------------ verify
say "verify"
HEALTH_BODY=""
for _ in $(seq "$HEALTH_WAIT"); do
if HEALTH_BODY="$(curl -fsS --max-time 2 "$HEALTH" 2>/dev/null)"; then break; fi
HEALTH_BODY=""
sleep 1
done
if [ -z "$HEALTH_BODY" ]; then
# 503 still means the daemon is up — it means herdr is unreachable. Say which.
CODE="$(curl -s -o /dev/null -w '%{http_code}' --max-time 2 "$HEALTH" 2>/dev/null || echo 000)"
if [ "$CODE" = "503" ]; then
warn "/healthz answers 503 degraded — the daemon is up but herdr is unreachable."
warn "Spawns will fail. Check herdr before delegating anything."
curl -s --max-time 2 "$HEALTH" 2>/dev/null | head -3 || true
else
die "/healthz never answered within ${HEALTH_WAIT}s (last code: $CODE). Last lines of $OUT:
$(tail -30 "$OUT" 2>/dev/null)"
fi
else
ok "/healthz 200 — $HEALTH_BODY"
warn "healthz green only proves herdr ANSWERS. If its protocol number changed, spawns can still"
warn "fail — prove a real spawn before trusting the fleet."
fi
# A fresh listening line, strictly after the restart mark. An old daemon that never died would
# otherwise let an old line pass for a new one.
if tail -n "+$((RESTART_MARK + 1))" "$OUT" 2>/dev/null | grep -q 'bridged listening'; then
ok "$(tail -n "+$((RESTART_MARK + 1))" "$OUT" | grep 'bridged listening' | tail -1)"
else
warn "no fresh 'bridged listening' line after the restart — check $OUT yourself"
fi
# Config keys the daemon accepted or deferred at boot. This is usually WHY you restarted.
say "config at boot"
tail -n "+$((RESTART_MARK + 1))" "$OUT" 2>/dev/null \
| grep -iE 'deferred|classification:|fleet health:|coverage' | tail -8 | sed 's/^/ /' \
|| echo " (nothing reported)"
# Errors since the restart, anchored to the marker so old noise cannot leak in.
ERRS="$(tail -n "+$((RESTART_MARK + 1))" "$OUT" 2>/dev/null | grep -cE ' (ERROR|SEVERE) ' || true)"
say "result"
ok "pid $NEW_PID, jar $(jar_id)"
if [ "${ERRS:-0}" -gt 0 ]; then
warn "$ERRS ERROR lines since restart:"
tail -n "+$((RESTART_MARK + 1))" "$OUT" | grep -E ' (ERROR|SEVERE) ' | tail -5 | sed 's/^/ /'
else
ok "no ERROR lines since restart"
fi
echo
echo " Next: call bridge_whoami and confirm it still answers 'primary'. A lead whose tab label"
echo " no longer matches fleet.leaders.*.tab is demoted to worker and refuses orchestration."
echo