fleetd #512 part 2: detect a died shutdown drain the ERROR count is blind to #534

Merged
ltms merged 1 commits from worker/512-part2-shutdown-detection-434701-9 into main 2026-09-12 08:04:15 +02:00
Member

Part 2 of fleetd #512. Part 1 (#522, the drain complete: released=N abandoned=M log line) is
already merged to main.

What changed

Scope: scripts/redeploy-fleetd.sh and scripts/test-redeploy-fleetd.sh only.

redeploy-fleetd.sh's existing ERROR-count classifier is structurally blind to the failure #493
is about: an uncaught exception in a shutdown thread goes straight to the JVM's default handler,
never through the logger, so it never carries an ERROR/SEVERE token — measured on two hosts,
including one where even a syslog priority filter cannot see it either.

Added:

  • scan_uncaught_exceptions — the negative check: greps the shutdown window for the failure's
    real shape (Exception in thread, NoClassDefFoundError), kept separate from
    classify_amqp_connection_errors (not an AMQP concern).
  • find_drain_complete_line — the positive check: looks for #522's
    drain complete: released=N abandoned=M (still BUSY at the shutdown deadline) line.
  • report_shutdown_drain — composes both into one decision+action function the main flow calls
    unconditionally (same shape as swap_if_built/refuse_drain_gate from #521/#528). Resolves to
    one of four outcomes:
    • complete — the drain-complete line is present: name the counts.
    • died — line absent, uncaught-exception shape found: the drain died. Name what was found.
    • unknown — line absent, no exception shape either: cannot tell, and said as such — never a
      pass or a failure. This is the ticket's trap: absence has two causes needing opposite handling
      (the previous daemon predates #522, or its drain failed without throwing), and this outcome is
      the third state that keeps them apart.
    • n/a — no previous daemon was actually stopped this run (cold start), so there is no shutdown
      window to have an opinion about.
  • Gated the "no ERROR lines since restart" summary line on the new outcome (item 4 of the
    ticket): it no longer prints when the drain died or the outcome is "cannot tell".

Never calls die() — warns loudly, does not fail the redeploy, per the ticket's explicit decision
(by the time this is detectable the new daemon is already up and healthy).

Tests

11 new test functions in test-redeploy-fleetd.sh (60 defined/invoked, was 49 on main):

  • both pure classifiers, present/absent fixtures
  • all four report_shutdown_drain outcomes as fixtures, including one with an uncaught exception
    and no line carrying an ERROR token (the heart of the ticket), a clean control, a
    "cannot tell" fixture, and an n/a fixture whose content deliberately looks like a died drain to
    prove the cold-start gate is actually consulted
  • a source-grep proof that the main flow's call site exists (sourcing stops before the main flow
    runs, so no behavioural test can see a deleted call), plus an ordering check
  • a source-grep proof that the "no ERROR lines" line is guarded by the new outcome

Verification (this worktree)

  • bash scripts/test-redeploy-fleetd.sh: exit 0. 60 test functions defined, 60 invoked (was
    49/49 on main). 0 lines matching ^FAIL: (the suite's own embedded mutation-proof tests print
    three lines containing "FAIL:" as expected internal output on a clean run — counted only
    ^FAIL:-anchored lines, which are the real failures).
  • bash -n scripts/redeploy-fleetd.sh and scripts/test-redeploy-fleetd.sh: pass under both
    /bin/bash (3.2.57) and env bash (5.3.9).
  • 5 mutations applied by hand and killed, each verified applied via two different greps plus a
    grep -n re-read, run against the suite, FAIL line and exit code quoted, restored, confirmed
    byte-identical via shasum -a 256, then a green control run:
    1. Inverted the died/unknown branch condition → FAIL: died fixture must set REDEPLOY_DRAIN_STATE=died: expected died, got unknown (exit 1).
    2. Inverted the had_previous_daemon gate → FAIL: died fixture must set REDEPLOY_DRAIN_STATE=died: expected died, got n/a (exit 1).
    3. Deleted the main flow's report_shutdown_drain call site → FAIL: could not find the main flow's report_shutdown_drain call site in redeploy-fleetd.sh (exit 1).
    4. Removed the REDEPLOY_DRAIN_STATE guard around "no ERROR lines since restart" → FAIL: 'no ERROR lines since restart' is not guarded by the shutdown-drain outcome (fleetd #512 item 4) (exit 1).
    5. Reintroduced the original defect (required an ERROR token in scan_uncaught_exceptions) →
      FAIL: scan must find the exception without an ERROR token: expected 1, got 0 (exit 1).
  • Did not run scripts/redeploy-fleetd.sh itself, with any flag, against the live daemon —
    tested only by sourcing it (as the suite already does) against fixture log files, per the
    ticket's hard constraint.

Caveats for review

  • Did not run mvn — this ticket is bash-only, no Java changed.
  • Could not run the CLAUDE.md wiki-sync check (wiki/ is uninitialized in this worktree by
    design) — not applicable here since neither file touched is part of the canonical
    CLAUDE.md block.
  • I added item 9 to the script's header comment (the "N remembered traps" list) describing this
    fix, matching the existing style for #492/#493/etc. Not explicitly requested by the ticket but
    consistent with how the file documents itself.
  • report_shutdown_drain's n/a outcome (no previous daemon stopped this run) is my own addition
    beyond the ticket's literal three-state wording — added because scanning a cold-start's own
    startup log for a previous daemon's shutdown drain would otherwise misreport "cannot tell" on
    every clean cold start. Flagging this so a reviewer can judge whether it's in scope; happy to
    simplify if not wanted.
Part 2 of fleetd #512. Part 1 (#522, the `drain complete: released=N abandoned=M` log line) is already merged to `main`. ## What changed Scope: `scripts/redeploy-fleetd.sh` and `scripts/test-redeploy-fleetd.sh` only. `redeploy-fleetd.sh`'s existing ERROR-count classifier is structurally blind to the failure #493 is about: an uncaught exception in a shutdown thread goes straight to the JVM's default handler, never through the logger, so it never carries an ERROR/SEVERE token — measured on two hosts, including one where even a syslog priority filter cannot see it either. Added: - `scan_uncaught_exceptions` — the negative check: greps the shutdown window for the failure's real shape (`Exception in thread`, `NoClassDefFoundError`), kept separate from `classify_amqp_connection_errors` (not an AMQP concern). - `find_drain_complete_line` — the positive check: looks for #522's `drain complete: released=N abandoned=M (still BUSY at the shutdown deadline)` line. - `report_shutdown_drain` — composes both into one decision+action function the main flow calls unconditionally (same shape as `swap_if_built`/`refuse_drain_gate` from #521/#528). Resolves to one of four outcomes: - `complete` — the drain-complete line is present: name the counts. - `died` — line absent, uncaught-exception shape found: the drain died. Name what was found. - `unknown` — line absent, no exception shape either: **cannot tell**, and said as such — never a pass or a failure. This is the ticket's trap: absence has two causes needing opposite handling (the previous daemon predates #522, or its drain failed without throwing), and this outcome is the third state that keeps them apart. - `n/a` — no previous daemon was actually stopped this run (cold start), so there is no shutdown window to have an opinion about. - Gated the `"no ERROR lines since restart"` summary line on the new outcome (item 4 of the ticket): it no longer prints when the drain died or the outcome is "cannot tell". Never calls `die()` — warns loudly, does not fail the redeploy, per the ticket's explicit decision (by the time this is detectable the new daemon is already up and healthy). ## Tests 11 new test functions in `test-redeploy-fleetd.sh` (60 defined/invoked, was 49 on `main`): - both pure classifiers, present/absent fixtures - all four `report_shutdown_drain` outcomes as fixtures, including one with an uncaught exception and **no** line carrying an ERROR token (the heart of the ticket), a clean control, a "cannot tell" fixture, and an `n/a` fixture whose content deliberately looks like a died drain to prove the cold-start gate is actually consulted - a source-grep proof that the main flow's call site exists (sourcing stops before the main flow runs, so no behavioural test can see a deleted call), plus an ordering check - a source-grep proof that the "no ERROR lines" line is guarded by the new outcome ## Verification (this worktree) - `bash scripts/test-redeploy-fleetd.sh`: exit 0. 60 test functions defined, 60 invoked (was 49/49 on `main`). 0 lines matching `^FAIL:` (the suite's own embedded mutation-proof tests print three lines *containing* "FAIL:" as expected internal output on a clean run — counted only `^FAIL:`-anchored lines, which are the real failures). - `bash -n scripts/redeploy-fleetd.sh` and `scripts/test-redeploy-fleetd.sh`: pass under both `/bin/bash` (3.2.57) and `env bash` (5.3.9). - 5 mutations applied by hand and killed, each verified applied via two different greps plus a `grep -n` re-read, run against the suite, FAIL line and exit code quoted, restored, confirmed byte-identical via `shasum -a 256`, then a green control run: 1. Inverted the died/unknown branch condition → `FAIL: died fixture must set REDEPLOY_DRAIN_STATE=died: expected died, got unknown` (exit 1). 2. Inverted the `had_previous_daemon` gate → `FAIL: died fixture must set REDEPLOY_DRAIN_STATE=died: expected died, got n/a` (exit 1). 3. Deleted the main flow's `report_shutdown_drain` call site → `FAIL: could not find the main flow's report_shutdown_drain call site in redeploy-fleetd.sh` (exit 1). 4. Removed the `REDEPLOY_DRAIN_STATE` guard around "no ERROR lines since restart" → `FAIL: 'no ERROR lines since restart' is not guarded by the shutdown-drain outcome (fleetd #512 item 4)` (exit 1). 5. Reintroduced the original defect (required an ERROR token in `scan_uncaught_exceptions`) → `FAIL: scan must find the exception without an ERROR token: expected 1, got 0` (exit 1). - Did **not** run `scripts/redeploy-fleetd.sh` itself, with any flag, against the live daemon — tested only by sourcing it (as the suite already does) against fixture log files, per the ticket's hard constraint. ## Caveats for review - Did not run `mvn` — this ticket is bash-only, no Java changed. - Could not run the CLAUDE.md wiki-sync check (`wiki/` is uninitialized in this worktree by design) — not applicable here since neither file touched is part of the canonical `CLAUDE.md` block. - I added item 9 to the script's header comment (the "N remembered traps" list) describing this fix, matching the existing style for #492/#493/etc. Not explicitly requested by the ticket but consistent with how the file documents itself. - `report_shutdown_drain`'s `n/a` outcome (no previous daemon stopped this run) is my own addition beyond the ticket's literal three-state wording — added because scanning a cold-start's own startup log for a *previous* daemon's shutdown drain would otherwise misreport "cannot tell" on every clean cold start. Flagging this so a reviewer can judge whether it's in scope; happy to simplify if not wanted.
agent added 1 commit 2026-09-12 07:56:38 +02:00
fleetd #512 part 2: detect a died shutdown drain the ERROR count is blind to
CI / contract (pull_request) Successful in 1m15s
CI / build (pull_request) Successful in 2m4s
190436c9cf
The previous daemon's dead shutdown drain (an uncaught exception in a
shutdown thread) never passes through the logger, so it never carries an
ERROR/SEVERE token, so redeploy-fleetd.sh's existing ERROR-count classifier
is structurally blind to it and prints a confident "no ERROR lines since
restart" while the drain actually died.

Add scan_uncaught_exceptions (greps the shutdown window for the failure's
real shape: `Exception in thread`, `NoClassDefFoundError`) and
find_drain_complete_line (checks for #522's SessionManager.drainAll
completion line). Compose both in report_shutdown_drain, a single
decision+action function the main flow calls unconditionally (same shape
as swap_if_built/refuse_drain_gate from #521/#528), which resolves to one
of four outcomes: complete, died, unknown ("cannot tell" — the line is
absent for either of two reasons that need opposite handling: the previous
daemon predates #522, or its drain failed without throwing), or n/a (no
previous daemon was actually stopped this run). Never fails the redeploy;
warns loudly instead.

Gate the "no ERROR lines since restart" summary line on the new outcome so
it never reads as reassurance when the drain died or the outcome is
"cannot tell" (item 4 of the ticket).

Tests: 11 new test functions (60 defined/invoked, was 49), covering both
pure classifiers, all four report_shutdown_drain outcomes, a source-grep
proof of the main-flow call site (sourcing stops before the main flow
runs), an ordering check, and the item-4 gating. Full suite green
(exit 0, 0 anchored FAIL lines). Five mutations applied and killed by hand
during review, each restored to a byte-identical file afterward.
ltms merged commit 7d711942fe into main 2026-09-12 08:04:15 +02:00
Sign in to join this conversation.