fleetd #511: fix wrong --no-build wording in drain-gate abort, pin jar_id() default #514

Merged
ltms merged 1 commits from worker/511-9a4b23-1 into main 2026-09-12 06:00:24 +02:00
Member

Fixes fleetd #511.

1 — drain-gate abort message (decision a, cheap honest fix)

The message told the operator a rerun "with or without --no-build" would finish the restart. That's wrong: by the time this message can fire, stage_built_jar has already moved the jar off $JAR, so a rerun with --no-build hits require_no_build_jar's own refusal ("no jar at $JAR — run without --no-build").

Before:

aborted — the running daemon was NOT touched, but the freshly built jar is sitting at
    $JAR_STAGED, not yet swapped into $JAR. Rerun (with or without --no-build) to finish the
    restart, or remove $JAR_STAGED by hand if you want to discard this build.

After:

aborted — the running daemon was NOT touched, but the freshly built jar is sitting at
    $JAR_STAGED, not yet swapped into $JAR. Rerun WITHOUT --no-build to finish the restart —
    the freshly built jar is no longer at the live path that --no-build requires — or
    remove $JAR_STAGED by hand if you want to discard this build.

Kept the staged path, the "daemon was NOT touched" fact, and the discard-by-hand option. Did not widen what --no-build means (that was option b, not taken).

Added test_drain_gate_abort_message_says_no_no_build, which reads the source (this code path is never reached by sourcing — the SOURCED guard stops before the main flow, same reason test_swap_ordered_after_wait_and_before_start reads source instead of calling a function) and asserts the old wrong phrase is gone and the corrected wording + reason are present.

2 — pin jar_id()'s default

Added test_jar_id_defaults_to_live_and_reports_explicit_path: sets JAR and JAR_STAGED to two files with different content, asserts a bare jar_id reports the hash of $JAR, and asserts jar_id "$JAR_STAGED" reports the other hash. Both halves, per the ticket.

Did not add a test for the rm -f "$JAR_STAGED" leftover wipe at line 578 — the ticket documents it as an equivalent mutant (mvn clean install deletes target/ on the next line regardless), so a test for it would only pin a line that can't change the outcome.

Verification

  • bash -n scripts/redeploy-fleetd.sh and bash -n scripts/test-redeploy-fleetd.sh: both pass.
  • bash scripts/test-redeploy-fleetd.sh: exit 0, 35 test functions defined and invoked (up from 33), same 3 internal FAIL: lines from the suite's own mutation cells, final PASS: redeploy log classifier.
  • Mutation proof for both new tests: reverted each fix, confirmed the mutation with two different greps + a grep -n re-read, ran the suite (went red with the expected FAIL: line, exit 1), restored from a saved copy, confirmed byte-identical with shasum -a 256, ran a green control (exit 0, same baseline shape).
  • Never ran scripts/redeploy-fleetd.sh against the live daemon; only sourced it inside the test suite as it already does.

Out of scope, noted per the ticket, not fixed: require_no_build_jar's own die message ("no jar at $JAR — run without --no-build") is itself an instance of the same "names a recovery" shape in miniature, but it's already correct and already pinned by test_require_no_build_jar_dies_when_absent — no new instance found beyond what's in the ticket.

Fixes fleetd #511. ## 1 — drain-gate abort message (decision a, cheap honest fix) The message told the operator a rerun "with or without `--no-build`" would finish the restart. That's wrong: by the time this message can fire, `stage_built_jar` has already moved the jar off `$JAR`, so a rerun with `--no-build` hits `require_no_build_jar`'s own refusal ("no jar at `$JAR` — run without `--no-build`"). Before: ``` aborted — the running daemon was NOT touched, but the freshly built jar is sitting at $JAR_STAGED, not yet swapped into $JAR. Rerun (with or without --no-build) to finish the restart, or remove $JAR_STAGED by hand if you want to discard this build. ``` After: ``` aborted — the running daemon was NOT touched, but the freshly built jar is sitting at $JAR_STAGED, not yet swapped into $JAR. Rerun WITHOUT --no-build to finish the restart — the freshly built jar is no longer at the live path that --no-build requires — or remove $JAR_STAGED by hand if you want to discard this build. ``` Kept the staged path, the "daemon was NOT touched" fact, and the discard-by-hand option. Did not widen what `--no-build` means (that was option b, not taken). Added `test_drain_gate_abort_message_says_no_no_build`, which reads the source (this code path is never reached by sourcing — the SOURCED guard stops before the main flow, same reason `test_swap_ordered_after_wait_and_before_start` reads source instead of calling a function) and asserts the old wrong phrase is gone and the corrected wording + reason are present. ## 2 — pin `jar_id()`'s default Added `test_jar_id_defaults_to_live_and_reports_explicit_path`: sets `JAR` and `JAR_STAGED` to two files with different content, asserts a bare `jar_id` reports the hash of `$JAR`, and asserts `jar_id "$JAR_STAGED"` reports the other hash. Both halves, per the ticket. Did **not** add a test for the `rm -f "$JAR_STAGED"` leftover wipe at line 578 — the ticket documents it as an equivalent mutant (`mvn clean install` deletes `target/` on the next line regardless), so a test for it would only pin a line that can't change the outcome. ## Verification - `bash -n scripts/redeploy-fleetd.sh` and `bash -n scripts/test-redeploy-fleetd.sh`: both pass. - `bash scripts/test-redeploy-fleetd.sh`: exit 0, 35 test functions defined and invoked (up from 33), same 3 internal `FAIL:` lines from the suite's own mutation cells, final `PASS: redeploy log classifier`. - Mutation proof for both new tests: reverted each fix, confirmed the mutation with two different greps + a `grep -n` re-read, ran the suite (went red with the expected `FAIL:` line, exit 1), restored from a saved copy, confirmed byte-identical with `shasum -a 256`, ran a green control (exit 0, same baseline shape). - Never ran `scripts/redeploy-fleetd.sh` against the live daemon; only sourced it inside the test suite as it already does. Out of scope, noted per the ticket, not fixed: `require_no_build_jar`'s own die message ("no jar at $JAR — run without --no-build") is itself an instance of the same "names a recovery" shape in miniature, but it's already correct and already pinned by `test_require_no_build_jar_dies_when_absent` — no new instance found beyond what's in the ticket.
agent added 1 commit 2026-09-12 05:57:39 +02:00
fleetd #511: fix wrong --no-build wording in drain-gate abort, pin jar_id() default
CI / contract (pull_request) Successful in 47s
CI / build (pull_request) Successful in 1m52s
6e23bf8309
The drain-gate abort message told the operator a rerun "with or without
--no-build" would finish the restart. That is wrong: by the time this
message can fire, stage_built_jar has already moved the jar off $JAR, so
--no-build hits require_no_build_jar's own refusal. Reworded to say the
rerun must NOT use --no-build, and why: the built jar is no longer at the
live path that --no-build requires.

Also added a test pinning jar_id()'s no-argument default (reports $JAR,
the live path) and its explicit-argument behavior (reports that path
instead), per fleetd #511 item 2. Not adding a test for the JAR_STAGED rm
-f at line 578 (fleetd #511 documents it as an equivalent mutant — mvn
clean install deletes target/ on the next line regardless).
ltms merged commit 525bc1c5f4 into main 2026-09-12 06:00:24 +02:00
Sign in to join this conversation.