fleetd #512 part 2: detect a died shutdown drain the ERROR count is blind to #534
Reference in New Issue
Block a user
Delete Branch "worker/512-part2-shutdown-detection-434701-9"
Deleting a branch is permanent. Although the deleted branch may continue to exist for a short time before it actually gets removed, it CANNOT be undone in most cases. Continue?
Part 2 of fleetd #512. Part 1 (#522, the
drain complete: released=N abandoned=Mlog line) isalready merged to
main.What changed
Scope:
scripts/redeploy-fleetd.shandscripts/test-redeploy-fleetd.shonly.redeploy-fleetd.sh's existing ERROR-count classifier is structurally blind to the failure #493is about: an uncaught exception in a shutdown thread goes straight to the JVM's default handler,
never through the logger, so it never carries an ERROR/SEVERE token — measured on two hosts,
including one where even a syslog priority filter cannot see it either.
Added:
scan_uncaught_exceptions— the negative check: greps the shutdown window for the failure'sreal shape (
Exception in thread,NoClassDefFoundError), kept separate fromclassify_amqp_connection_errors(not an AMQP concern).find_drain_complete_line— the positive check: looks for #522'sdrain complete: released=N abandoned=M (still BUSY at the shutdown deadline)line.report_shutdown_drain— composes both into one decision+action function the main flow callsunconditionally (same shape as
swap_if_built/refuse_drain_gatefrom #521/#528). Resolves toone of four outcomes:
complete— the drain-complete line is present: name the counts.died— line absent, uncaught-exception shape found: the drain died. Name what was found.unknown— line absent, no exception shape either: cannot tell, and said as such — never apass or a failure. This is the ticket's trap: absence has two causes needing opposite handling
(the previous daemon predates #522, or its drain failed without throwing), and this outcome is
the third state that keeps them apart.
n/a— no previous daemon was actually stopped this run (cold start), so there is no shutdownwindow to have an opinion about.
"no ERROR lines since restart"summary line on the new outcome (item 4 of theticket): it no longer prints when the drain died or the outcome is "cannot tell".
Never calls
die()— warns loudly, does not fail the redeploy, per the ticket's explicit decision(by the time this is detectable the new daemon is already up and healthy).
Tests
11 new test functions in
test-redeploy-fleetd.sh(60 defined/invoked, was 49 onmain):report_shutdown_drainoutcomes as fixtures, including one with an uncaught exceptionand no line carrying an ERROR token (the heart of the ticket), a clean control, a
"cannot tell" fixture, and an
n/afixture whose content deliberately looks like a died drain toprove the cold-start gate is actually consulted
runs, so no behavioural test can see a deleted call), plus an ordering check
Verification (this worktree)
bash scripts/test-redeploy-fleetd.sh: exit 0. 60 test functions defined, 60 invoked (was49/49 on
main). 0 lines matching^FAIL:(the suite's own embedded mutation-proof tests printthree lines containing "FAIL:" as expected internal output on a clean run — counted only
^FAIL:-anchored lines, which are the real failures).bash -n scripts/redeploy-fleetd.shandscripts/test-redeploy-fleetd.sh: pass under both/bin/bash(3.2.57) andenv bash(5.3.9).grep -nre-read, run against the suite, FAIL line and exit code quoted, restored, confirmedbyte-identical via
shasum -a 256, then a green control run:FAIL: died fixture must set REDEPLOY_DRAIN_STATE=died: expected died, got unknown(exit 1).had_previous_daemongate →FAIL: died fixture must set REDEPLOY_DRAIN_STATE=died: expected died, got n/a(exit 1).report_shutdown_draincall site →FAIL: could not find the main flow's report_shutdown_drain call site in redeploy-fleetd.sh(exit 1).REDEPLOY_DRAIN_STATEguard around "no ERROR lines since restart" →FAIL: 'no ERROR lines since restart' is not guarded by the shutdown-drain outcome (fleetd #512 item 4)(exit 1).scan_uncaught_exceptions) →FAIL: scan must find the exception without an ERROR token: expected 1, got 0(exit 1).scripts/redeploy-fleetd.shitself, with any flag, against the live daemon —tested only by sourcing it (as the suite already does) against fixture log files, per the
ticket's hard constraint.
Caveats for review
mvn— this ticket is bash-only, no Java changed.wiki/is uninitialized in this worktree bydesign) — not applicable here since neither file touched is part of the canonical
CLAUDE.mdblock.fix, matching the existing style for #492/#493/etc. Not explicitly requested by the ticket but
consistent with how the file documents itself.
report_shutdown_drain'sn/aoutcome (no previous daemon stopped this run) is my own additionbeyond the ticket's literal three-state wording — added because scanning a cold-start's own
startup log for a previous daemon's shutdown drain would otherwise misreport "cannot tell" on
every clean cold start. Flagging this so a reviewer can judge whether it's in scope; happy to
simplify if not wanted.